EveryDev.ai
Subscribe
Home
Tools

4,178+ AI tools

  • New
  • Trending
  • Featured
  • Rate tools
  • Compare
  • Arena
Categories
  • Agents3186
  • Coding2210
  • Infrastructure966
  • Projects660
  • Marketing625
  • Research579
  • MCP518
  • Design497
  • Analytics490
  • Testing382
  • Security360
  • Data320
  • Integration242
  • Prompts239
  • Communication225
  • Extensions210
  • Voice191
  • Learning188
  • Commerce167
  • DevOps145
  • Web99
  • Finance34
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. ReviewBench
    R

    ReviewBench

    LLM Evaluations

    An open, reproducible benchmark and leaderboard for evaluating AI code review agents on real-world pull requests.

    Visit Website

    At a Glance

    Pricing
    Open Source

    MIT-licensed benchmark artifacts. Running your reviewer consumes your own model API usage; portal evaluation costs and terms follow the official README.

    Engagement

    Available On

    Web
    CLI
    Linux
    macOS
    Windows

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsCode ReviewPerformance Metrics

    Alternatives

    OmnisBenchNeedle In A HaystackCompute:Arena
    Developer
    GitHubSan FranciscoEst. 2008$350M raised

    Listed Oct 2026

    About ReviewBench

    ReviewBench is an open benchmark, currently a research preview, developed by GitHub Inc. for measuring how well AI code review agents find issues in real pull requests. Agents are scored against a human-validated reference set of findings, and results are published on a public leaderboard. It is a benchmark rather than an AI reviewer itself.

    What It Is

    ReviewBench evaluates AI code review systems. The benchmark has 219 public pull requests from 187 repositories across 19 languages. According to GitHub's announcement, language and repository-size distributions were informed by an analysis of 103.9 million GitHub pull requests, with pull request size weighted toward more substantive changes. Each pull request has a golden set of findings gathered from human reviewers, author follow-up commits, static analysis and multiple frontier LLMs. Findings are deduplicated and validated under a shared rubric, and each is labeled by severity and category.

    How Scoring Works

    The benchmark reports metrics in two families. Grounded precision, recall and F1 use only the existing golden-set labels and give the strict, like-for-like comparison. Augmented metrics also count findings outside the golden set, which an LLM judge labels as real or not, so agents get credit for new discoveries. Because augmented recall depends on each agent's own discoveries, it is not directly comparable across agents. The leaderboard ranks by grounded F1 and can be re-sliced by severity, category and precision or recall preference (F0.5 or F2). GitHub reports that senior engineers' independent labels agreed with the benchmark's true-positive labels 96.6% of the time.

    Running and Submitting an Agent

    Users wrap an agent in a container that follows the published agent contract, then validate it locally on a 25-PR test set with a provided script. Agents are registered through the website after signing in with GitHub, using a container image, configuration and the submitter's own model key. A final run covers all 219 pull requests over three rounds and is scored by the benchmark judge. A maintainer approves results before they appear on the leaderboard. The dataset, judge prompts and a local judging CLI are public, so teams can also score findings privately while tuning.

    Context and Caveats

    The site states that the initial leaderboard entries were produced by the ReviewBench team running each vendor's public product, and that vendors did not verify them. GitHub also makes Copilot code review, one of the evaluated products. The site says results are for research and informational purposes, rely partly on AI-assisted judgments, and may not predict performance on a user's own code.

    ReviewBench - 1

    Community Discussions

    Be the first to start a conversation about ReviewBench

    Share your experience with ReviewBench, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    MIT-licensed benchmark artifacts. Running your reviewer consumes your own model API usage; portal evaluation costs and terms follow the official README.

    • MIT License
    • Full benchmark set: 219 pull requests with human-reviewed golden findings
    • Test set of 25 tasks
    • Public self-serve runner and judging CLI
    • Bring your own container image and model key

    Capabilities

    Key Features

    • 219 pull requests from 187 repositories across 19 languages
    • Golden set of findings from human reviewers, static analysis and frontier LLMs
    • Severity and category labels for findings
    • Grounded and augmented precision, recall and F1 metrics
    • Adjustable precision/recall preference (F0.5, F1, F2) filtering
    • Public leaderboard with agent comparison
    • Bring-your-own container image and model key
    • 25-PR test set and local try-agent script
    • Local judging CLI for private tuning
    • Published methodology, rubric and judge prompts

    Integrations

    GitHub
    Docker
    GitHub Container Registry
    Codex CLI

    Ratings & Reviews

    No ratings yet

    Be the first to rate ReviewBench and help others make informed decisions.

    Rate other tools you’ve used

    Developer

    GitHub

    GitHub is a Microsoft-owned platform for software development and version control using Git, offering tools like Copilot, CLI, and collaborative development features.

    Founded 2008
    88 Colin P Kelly Jr St, CA 94107
    $350M raised
    5,000 employees

    Used by

    Ford
    Stripe
    PagerDuty
    Pinterest
    +4 more
    Read more about GitHub
    WebsiteGitHubX / Twitter
    12 tools in directory

    Similar Tools

    OmnisBench icon

    OmnisBench

    An open, reproducible benchmark for LLM routing efficiency that measures how close a routing policy gets to the ideal quality-per-dollar frontier using fresh, uncontaminated tasks.

    Needle In A Haystack icon

    Needle In A Haystack

    A CLI tool that pressure-tests LLM long-context retrieval by sweeping context length and needle depth combinations to measure model accuracy.

    Compute:Arena icon

    Compute:Arena

    Community-driven benchmark leaderboard for local AI models on edge hardware, with a CLI tool to run, sign, and submit throughput results.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    135 tools

    Code Review

    Tools that help review, analyze, and improve code quality.

    122 tools

    Performance Metrics

    Specialized tools for measuring, evaluating, and optimizing AI model performance across accuracy, speed, resource utilization, and other critical parameters.

    65 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions