# ReviewBench

> An open, reproducible benchmark and leaderboard for evaluating AI code review agents on real-world pull requests.

ReviewBench is an open benchmark, currently a research preview, developed by GitHub Inc. for measuring how well AI code review agents find issues in real pull requests. Agents are scored against a human-validated reference set of findings, and results are published on a public leaderboard. It is a benchmark rather than an AI reviewer itself.

## What It Is

ReviewBench evaluates AI code review systems. The benchmark has 219 public pull requests from 187 repositories across 19 languages. According to GitHub's announcement, language and repository-size distributions were informed by an analysis of 103.9 million GitHub pull requests, with pull request size weighted toward more substantive changes. Each pull request has a golden set of findings gathered from human reviewers, author follow-up commits, static analysis and multiple frontier LLMs. Findings are deduplicated and validated under a shared rubric, and each is labeled by severity and category.

## How Scoring Works

The benchmark reports metrics in two families. Grounded precision, recall and F1 use only the existing golden-set labels and give the strict, like-for-like comparison. Augmented metrics also count findings outside the golden set, which an LLM judge labels as real or not, so agents get credit for new discoveries. Because augmented recall depends on each agent's own discoveries, it is not directly comparable across agents. The leaderboard ranks by grounded F1 and can be re-sliced by severity, category and precision or recall preference (F0.5 or F2). GitHub reports that senior engineers' independent labels agreed with the benchmark's true-positive labels 96.6% of the time.

## Running and Submitting an Agent

Users wrap an agent in a container that follows the published agent contract, then validate it locally on a 25-PR test set with a provided script. Agents are registered through the website after signing in with GitHub, using a container image, configuration and the submitter's own model key. A final run covers all 219 pull requests over three rounds and is scored by the benchmark judge. A maintainer approves results before they appear on the leaderboard. The dataset, judge prompts and a local judging CLI are public, so teams can also score findings privately while tuning.

## Context and Caveats

The site states that the initial leaderboard entries were produced by the ReviewBench team running each vendor's public product, and that vendors did not verify them. GitHub also makes Copilot code review, one of the evaluated products. The site says results are for research and informational purposes, rely partly on AI-assisted judgments, and may not predict performance on a user's own code.

## Features
- 219 pull requests from 187 repositories across 19 languages
- Golden set of findings from human reviewers, static analysis and frontier LLMs
- Severity and category labels for findings
- Grounded and augmented precision, recall and F1 metrics
- Adjustable precision/recall preference (F0.5, F1, F2) filtering
- Public leaderboard with agent comparison
- Bring-your-own container image and model key
- 25-PR test set and local try-agent script
- Local judging CLI for private tuning
- Published methodology, rubric and judge prompts

## Integrations
GitHub, Docker, GitHub Container Registry, Codex CLI

## Platforms
WEB, CLI, LINUX, MACOS, WINDOWS

## Pricing
Open Source

## Links
- Website: https://review-bench.ai/
- Documentation: https://github.com/review-bench/ReviewBench/blob/main/docs/ONBOARDING.md
- Repository: https://github.com/review-bench/ReviewBench
- EveryDev.ai: https://www.everydev.ai/tools/reviewbench
