GitHub launches ReviewBench for AI code review evaluation
The benchmark is meant to make offline comparisons between code reviewers more useful for production decisions.
GitHub has introduced ReviewBench, an offline benchmark for AI code review, and says it is meant to make comparisons between reviewers more rigorous and more useful for production decisions. The benchmark was built from 103.9 million GitHub pull requests and contains 219 pull requests from 187 public open source repositories across 19 languages, with distributions chosen to resemble GitHub’s overall workload. GitHub says it assembles ground truth from human reviewers, follow-up commits, analysis tools, and multiple LLMs, then deduplicates and grades findings with a shared rubric using Claude Sonnet 5. It also reports both grounded and augmented precision/recall/F1 so systems can be judged on known issues as well as valid discoveries outside the initial gold set. GitHub says ReviewBench is publicly available and that it already improved how well Copilot code review offline results predict production experiment outcomes.
Why it matters
For teams assessing AI code review tools, ReviewBench gives a more structured way to compare systems against the kinds of issues that matter in practice. GitHub says it also makes offline evaluation better at predicting production experiments for Copilot code review, which could narrow the gap between lab results and real-world behavior.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- GitHub Blog