ReviewBench is GitHub's new offline exam for AI code review agents. On October 5, 2026, GitHub published the open benchmark on its official blog: after analyzing the distribution of 103.9 million real pull requests, it selected 219 public PRs from 187 open-source repositories across 19 languages, so different code review agents can be compared under one golden set and one scoring rubric — who misses less, who cries wolf less, and who catches the critical problems. ReviewBench is in research preview, the full dataset is public, and any team can run its own review agent against it and submit results.
The golden set was not decided by one reviewer
Code review is hard to evaluate because the ground truth itself is contested: on the same PR, human reviewers, static analysis tools, and different large models each find a different set of issues, and leaving out any one perspective skews the exam toward one kind of reviewer. ReviewBench builds its golden set from multiple sources: candidate findings come from real human review comments, issues inferred from authors' follow-up commits, deterministic analysis tools, and several frontier LLMs. Overlapping findings are semantically deduplicated, then judged under one published rubric — a finding counts only if it is true, relevant, and non-trivial, regardless of which source produced it. Claude Sonnet 5 serves as the judge, with the rubric and judge configuration published. Every finding is labeled by severity (critical, medium, low) and category (correctness, security, reliability, maintainability, testing, and more). Before release, senior engineers who had not built the dataset independently re-labeled every golden finding, agreeing with the benchmark 96.6% of the time.
New metrics leave room for discoveries outside the answer key
A conventional benchmark scores precision and recall against a fixed answer key, so a reviewer that finds a genuine problem missing from the key gets punished for it. ReviewBench therefore reports two families of metrics: one grounded strictly in the golden set for apples-to-apples comparison, and an augmented set in which the judge independently decides whether unmatched new findings are true or false, crediting real discoveries. Results can also be sliced by severity and category, and the Fβ weight can be adjusted between "miss nothing, especially not critical issues" and "keep the noise down," re-ranking the leaderboard accordingly — because for code review, no single preference fits every team.
Can offline scores predict production? GitHub tested it on itself
GitHub first used the benchmark internally to iterate on Copilot code review, and the post gives a complete case study: in a multi-model ensemble review experiment, ReviewBench predicted higher precision, recall, and comment volume at lower cost per review. The subsequent online A/B test moved the same way — the share of comments developers acted on rose 8.0%, recall rose 13.6%, comment volume rose 61%, and cost per review fell 8.0%. For critical comments, the offline prediction was a 227% increase against 262% online, with even the severity distribution lining up.
The submission path for outside teams is open: sign in with a GitHub account on the ReviewBench site, register the agent's container image, configuration, and your own model key, tune on a 25-PR test set, then run the full set of 219 PRs over three rounds; scores reach the leaderboard only after maintainer review, and only replace an existing entry when they improve on it. AI code review is becoming a default part of the development workflow — from Codex's rapid iteration to a crowd of review bots — yet what the field lacked was a public, shared yardstick that also distinguishes severity. ReviewBench publishes the rubric, the dataset, and the exam itself; the question now is which mainstream review tools will dare to post their scores.