EveryDev.ai
Subscribe
Home
Tools

4,304+ AI tools

  • New
  • Trending
  • Featured
  • Rate tools
  • Compare
  • Arena
Categories
  • Agents3274
  • Coding2275
  • Infrastructure1000
  • Projects696
  • Marketing636
  • Research587
  • MCP532
  • Design508
  • Analytics506
  • Testing394
  • Security376
  • Data327
  • Integration244
  • Prompts244
  • Communication235
  • Extensions217
  • Voice193
  • Learning190
  • Commerce170
  • DevOps153
  • Web103
  • Finance36
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. Benchmark Heaven
    Benchmark Heaven icon

    Benchmark Heaven

    LLM Evaluations
    Featured

    Open-source web tool that compares AI model benchmark results with modeled per-task costs across providers.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Benchmark Heaven code is open source under the MIT licence; the licence covers the repository's code only, not the third-party benchmark and price data it collects.

    Engagement

    Available On

    Web
    API

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsModel ManagementPerformance Metrics

    Alternatives

    OmnisBenchNeedle In A HaystackReviewBench
    Developer
    Florian StandhartingerFlorian Standhartinger operates productivity-boost.com Betri…

    Listed Oct 2026

    About Benchmark Heaven

    Benchmark Heaven is a free, open-source hobby project that brings benchmark results and provider pricing for open-source and frontier LLMs into one comparable view. It is built by Florian Standhartinger and is labelled as beta, with data and features changing daily. The site states it tracks 895 models, 284 benchmarks and 19,875 results, with a dataset dated 2026-10-09.

    What It Is

    Benchmark Heaven is a model comparison and leaderboard site. It combines capability data from Artificial Analysis, Epoch AI and DesignArena with prices from OpenRouter, AWS Bedrock, Azure AI Foundry, Google Vertex AI and European providers, normalized to USD per 1M tokens. Users can rank models, compare them side by side, and see which models give the most capability at a given price.

    How the Scoring and Cost Work

    The Main Composite score is rank-based: each result becomes a percentile on a common scale, and seven slots (AA Coding, AA Coding Agent v1.4, AA Intelligence, Epoch ECI, Epoch Software ECI, DesignArena Web Apps and Full-Stack) are averaged. Missing slots are imputed from the model's own mean, and thinly measured models carry a Thin data badge. Adjusted cost estimates the USD cost of one task from the cheapest eligible provider route, its cache behaviour and the model's measured token usage on a common input/output workload. A value map plots capability against cost with a Pareto frontier line.

    Benchmaxxing Signal

    The site flags models that rank higher on famous public benchmarks than on held-out ones, as a screening signal rather than proof of intent. It can optionally be included in the score.

    Views and Filters

    Pages cover benchmark tables, compare, charts, cost vs capability, providers per model, a provider explorer, gateways, and EU and sovereign hosting. Filters include region of hosting, provider data policy, open weights only and deprecated models. A subscription estimator compares plan costs with API costs. JevBench, AudioJevBench and ImageJevBench are the site's own decision-model benchmarks. A public read-only JSON API and WebMCP tools in the browser are also provided.

    Licensing and Data

    The code is MIT-licensed. Collected benchmark and price data remains third-party data under each source's own terms.

    Benchmark Heaven - 1

    Community Discussions

    Be the first to start a conversation about Benchmark Heaven

    Share your experience with Benchmark Heaven, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source (MIT)

    Benchmark Heaven code is open source under the MIT licence; the licence covers the repository's code only, not the third-party benchmark and price data it collects.

    • MIT-licensed code (repository code only)
    • Third-party benchmark results, prices and other data are not relicensed; each source keeps its own terms
    • Public read-only JSON API (CORS-enabled)

    Capabilities

    Key Features

    • Benchmark results for hundreds of models with sources and dates
    • Rank-based Main Composite capability score
    • Modeled adjusted cost per task
    • Cost vs capability scatter with Pareto frontier
    • Side-by-side model comparison
    • Benchmaxxing signal
    • Provider and gateway explorers
    • EU and sovereign hosting filters
    • Subscription plan cost estimator
    • JevBench decision model benchmarks
    • Public read-only JSON API
    • WebMCP tools for agents

    Integrations

    Artificial Analysis
    Epoch AI
    DesignArena
    OpenRouter
    AWS Bedrock
    Azure AI Foundry
    Google Vertex AI
    GitHub Copilot
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate Benchmark Heaven and help others make informed decisions.

    Rate other tools you’ve used

    Developer

    Florian Standhartinger

    Florian Standhartinger operates productivity-boost.com Betriebs UG (haftungsbeschränkt) & Co. KG, a one-person company. He builds Benchmark Heaven as an open-source hobby project to help people choose which LLM fits a job.

    Read more about Florian Standhartinger
    GitHub
    1 tool in directory

    Similar Tools

    OmnisBench icon

    OmnisBench

    An open, reproducible benchmark for LLM routing efficiency that measures how close a routing policy gets to the ideal quality-per-dollar frontier using fresh, uncontaminated tasks.

    Needle In A Haystack icon

    Needle In A Haystack

    A CLI tool that pressure-tests LLM long-context retrieval by sweeping context length and needle depth combinations to measure model accuracy.

    ReviewBench icon

    ReviewBench

    An open, reproducible benchmark and leaderboard for evaluating AI code review agents on real-world pull requests.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    141 tools

    Model Management

    Tools for managing, versioning, and deploying AI models.

    70 tools

    Performance Metrics

    Specialized tools for measuring, evaluating, and optimizing AI model performance across accuracy, speed, resource utilization, and other critical parameters.

    69 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions
    1upvote