Benchmark Heaven
Open-source web tool that compares AI model benchmark results with modeled per-task costs across providers.
At a Glance
Benchmark Heaven code is open source under the MIT licence; the licence covers the repository's code only, not the third-party benchmark and price data it collects.
Engagement
Available On
Alternatives
Listed Oct 2026
About Benchmark Heaven
Benchmark Heaven is a free, open-source hobby project that brings benchmark results and provider pricing for open-source and frontier LLMs into one comparable view. It is built by Florian Standhartinger and is labelled as beta, with data and features changing daily. The site states it tracks 895 models, 284 benchmarks and 19,875 results, with a dataset dated 2026-10-09.
What It Is
Benchmark Heaven is a model comparison and leaderboard site. It combines capability data from Artificial Analysis, Epoch AI and DesignArena with prices from OpenRouter, AWS Bedrock, Azure AI Foundry, Google Vertex AI and European providers, normalized to USD per 1M tokens. Users can rank models, compare them side by side, and see which models give the most capability at a given price.
How the Scoring and Cost Work
The Main Composite score is rank-based: each result becomes a percentile on a common scale, and seven slots (AA Coding, AA Coding Agent v1.4, AA Intelligence, Epoch ECI, Epoch Software ECI, DesignArena Web Apps and Full-Stack) are averaged. Missing slots are imputed from the model's own mean, and thinly measured models carry a Thin data badge. Adjusted cost estimates the USD cost of one task from the cheapest eligible provider route, its cache behaviour and the model's measured token usage on a common input/output workload. A value map plots capability against cost with a Pareto frontier line.
Benchmaxxing Signal
The site flags models that rank higher on famous public benchmarks than on held-out ones, as a screening signal rather than proof of intent. It can optionally be included in the score.
Views and Filters
Pages cover benchmark tables, compare, charts, cost vs capability, providers per model, a provider explorer, gateways, and EU and sovereign hosting. Filters include region of hosting, provider data policy, open weights only and deprecated models. A subscription estimator compares plan costs with API costs. JevBench, AudioJevBench and ImageJevBench are the site's own decision-model benchmarks. A public read-only JSON API and WebMCP tools in the browser are also provided.
Licensing and Data
The code is MIT-licensed. Collected benchmark and price data remains third-party data under each source's own terms.
Community Discussions
Be the first to start a conversation about Benchmark Heaven
Share your experience with Benchmark Heaven, ask questions, or help others learn from your insights.
Pricing
Open Source (MIT)
Benchmark Heaven code is open source under the MIT licence; the licence covers the repository's code only, not the third-party benchmark and price data it collects.
- MIT-licensed code (repository code only)
- Third-party benchmark results, prices and other data are not relicensed; each source keeps its own terms
- Public read-only JSON API (CORS-enabled)
Capabilities
Key Features
- Benchmark results for hundreds of models with sources and dates
- Rank-based Main Composite capability score
- Modeled adjusted cost per task
- Cost vs capability scatter with Pareto frontier
- Side-by-side model comparison
- Benchmaxxing signal
- Provider and gateway explorers
- EU and sovereign hosting filters
- Subscription plan cost estimator
- JevBench decision model benchmarks
- Public read-only JSON API
- WebMCP tools for agents
