ReviewBench
An open, reproducible benchmark and leaderboard for evaluating AI code review agents on real-world pull requests.
At a Glance
MIT-licensed benchmark artifacts. Running your reviewer consumes your own model API usage; portal evaluation costs and terms follow the official README.
Engagement
Available On
Alternatives
Listed Oct 2026
About ReviewBench
ReviewBench is an open benchmark, currently a research preview, developed by GitHub Inc. for measuring how well AI code review agents find issues in real pull requests. Agents are scored against a human-validated reference set of findings, and results are published on a public leaderboard. It is a benchmark rather than an AI reviewer itself.
What It Is
ReviewBench evaluates AI code review systems. The benchmark has 219 public pull requests from 187 repositories across 19 languages. According to GitHub's announcement, language and repository-size distributions were informed by an analysis of 103.9 million GitHub pull requests, with pull request size weighted toward more substantive changes. Each pull request has a golden set of findings gathered from human reviewers, author follow-up commits, static analysis and multiple frontier LLMs. Findings are deduplicated and validated under a shared rubric, and each is labeled by severity and category.
How Scoring Works
The benchmark reports metrics in two families. Grounded precision, recall and F1 use only the existing golden-set labels and give the strict, like-for-like comparison. Augmented metrics also count findings outside the golden set, which an LLM judge labels as real or not, so agents get credit for new discoveries. Because augmented recall depends on each agent's own discoveries, it is not directly comparable across agents. The leaderboard ranks by grounded F1 and can be re-sliced by severity, category and precision or recall preference (F0.5 or F2). GitHub reports that senior engineers' independent labels agreed with the benchmark's true-positive labels 96.6% of the time.
Running and Submitting an Agent
Users wrap an agent in a container that follows the published agent contract, then validate it locally on a 25-PR test set with a provided script. Agents are registered through the website after signing in with GitHub, using a container image, configuration and the submitter's own model key. A final run covers all 219 pull requests over three rounds and is scored by the benchmark judge. A maintainer approves results before they appear on the leaderboard. The dataset, judge prompts and a local judging CLI are public, so teams can also score findings privately while tuning.
Context and Caveats
The site states that the initial leaderboard entries were produced by the ReviewBench team running each vendor's public product, and that vendors did not verify them. GitHub also makes Copilot code review, one of the evaluated products. The site says results are for research and informational purposes, rely partly on AI-assisted judgments, and may not predict performance on a user's own code.
Community Discussions
Be the first to start a conversation about ReviewBench
Share your experience with ReviewBench, ask questions, or help others learn from your insights.
Pricing
Open Source
MIT-licensed benchmark artifacts. Running your reviewer consumes your own model API usage; portal evaluation costs and terms follow the official README.
- MIT License
- Full benchmark set: 219 pull requests with human-reviewed golden findings
- Test set of 25 tasks
- Public self-serve runner and judging CLI
- Bring your own container image and model key
Capabilities
Key Features
- 219 pull requests from 187 repositories across 19 languages
- Golden set of findings from human reviewers, static analysis and frontier LLMs
- Severity and category labels for findings
- Grounded and augmented precision, recall and F1 metrics
- Adjustable precision/recall preference (F0.5, F1, F2) filtering
- Public leaderboard with agent comparison
- Bring-your-own container image and model key
- 25-PR test set and local try-agent script
- Local judging CLI for private tuning
- Published methodology, rubric and judge prompts
