# Benchmark Heaven

> Open-source web tool that compares AI model benchmark results with modeled per-task costs across providers.

Benchmark Heaven is a free, open-source hobby project that brings benchmark results and provider pricing for open-source and frontier LLMs into one comparable view. It is built by Florian Standhartinger and is labelled as beta, with data and features changing daily. The site states it tracks 895 models, 284 benchmarks and 19,875 results, with a dataset dated 2026-10-09.

## What It Is

Benchmark Heaven is a model comparison and leaderboard site. It combines capability data from Artificial Analysis, Epoch AI and DesignArena with prices from OpenRouter, AWS Bedrock, Azure AI Foundry, Google Vertex AI and European providers, normalized to USD per 1M tokens. Users can rank models, compare them side by side, and see which models give the most capability at a given price.

## How the Scoring and Cost Work

The Main Composite score is rank-based: each result becomes a percentile on a common scale, and seven slots (AA Coding, AA Coding Agent v1.4, AA Intelligence, Epoch ECI, Epoch Software ECI, DesignArena Web Apps and Full-Stack) are averaged. Missing slots are imputed from the model's own mean, and thinly measured models carry a Thin data badge. Adjusted cost estimates the USD cost of one task from the cheapest eligible provider route, its cache behaviour and the model's measured token usage on a common input/output workload. A value map plots capability against cost with a Pareto frontier line.

## Benchmaxxing Signal

The site flags models that rank higher on famous public benchmarks than on held-out ones, as a screening signal rather than proof of intent. It can optionally be included in the score.

## Views and Filters

Pages cover benchmark tables, compare, charts, cost vs capability, providers per model, a provider explorer, gateways, and EU and sovereign hosting. Filters include region of hosting, provider data policy, open weights only and deprecated models. A subscription estimator compares plan costs with API costs. JevBench, AudioJevBench and ImageJevBench are the site's own decision-model benchmarks. A public read-only JSON API and WebMCP tools in the browser are also provided.

## Licensing and Data

The code is MIT-licensed. Collected benchmark and price data remains third-party data under each source's own terms.

## Features
- Benchmark results for hundreds of models with sources and dates
- Rank-based Main Composite capability score
- Modeled adjusted cost per task
- Cost vs capability scatter with Pareto frontier
- Side-by-side model comparison
- Benchmaxxing signal
- Provider and gateway explorers
- EU and sovereign hosting filters
- Subscription plan cost estimator
- JevBench decision model benchmarks
- Public read-only JSON API
- WebMCP tools for agents

## Integrations
Artificial Analysis, Epoch AI, DesignArena, OpenRouter, AWS Bedrock, Azure AI Foundry, Google Vertex AI, GitHub Copilot

## Platforms
WEB, API

## Pricing
Open Source

## Links
- Website: https://benchmarkheaven.com
- Documentation: https://benchmarkheaven.com/about
- Repository: https://github.com/fstandhartinger/model-market-comparison
- EveryDev.ai: https://www.everydev.ai/tools/benchmark-heaven
