Zero Day Room
Live
Vulnerabilities

GitHub and Microsoft launch ReviewBench AI code review benchmark

GitHub and Microsoft have released ReviewBench, an open benchmark for evaluating AI code review tools using 219 real-world pull requests.

GitHub and Microsoft have released ReviewBench, an open benchmark for evaluating AI code review tools using 219...

GitHub and Microsoft have launched ReviewBench, an open benchmark for evaluating AI code review tools. It tests agents on 219 pull requests from 187 public repositories across 19 programming languages, using a rigorous, multi-source ground truth process.

ReviewBench reports six core metrics. Grounded metrics measure performance against a fixed set of known issues, while augmented metrics account for valid findings outside that set. The benchmark is available in research preview and was built by a joint team from GitHub and Microsoft. It allows results to be sliced by issue severity and category, enabling detailed analysis.

The benchmark uses a rigorous, multi-source ground truth process

GitHub built a reference collection, called the golden set, in three stages. They gathered candidate findings from human reviewers, follow-up code changes, static analysis tools, and AI models. Findings describing the same issue were merged. Each was then assessed using a shared evaluation rubric for reliability.

The benchmark's grounded metrics measure precision-the share of an agent's findings that match known issues-and recall, which is the proportion of known issues detected. The F1 score gives equal weight to both. For augmented metrics, an AI judge assesses whether findings outside the golden set are valid. GitHub uses grounded recall as the main metric for comparing different systems side-by-side. To validate the process, senior engineers independently re-labeled every ground-truth finding. Their judgments agreed with ReviewBench's 96.6% of the time.

ReviewBench enables flexible, reproducible evaluation for teams

Users can tailor evaluations to their needs. They can adjust the beta parameter in the F-beta score to weight recall for broader coverage or precision for lower noise. Results can be filtered by severity-critical, medium, or low-and by category, such as security, correctness, or maintainability. The leaderboard updates rankings based on these user preferences.

To submit an agent, users sign in with GitHub and provide a container image, configuration, and model access key. A 25-pull-request test set allows for performance assessment and refinement. Three rounds are then run on the full benchmark. ReviewBench scores every agent with the same AI judge for consistency. Scores remain private until a maintainer approves the submission. They are published only for an agent's first leaderboard entry or when it exceeds a previous score.

Pull request size in ReviewBench is weighted. Michelle Zhou, Data Scientist at Microsoft, and Alejandro Carderera de Diego, Senior Applied Researcher at GitHub said: "This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most."

GitHub used ReviewBench to guide improvements in Copilot Code Review

The benchmark provided an early offline signal that predicted the direction of production experiments for GitHub's own Copilot Code Review (CCR). GitHub used it to evaluate successive iterations before user testing. In one experiment with the lite tier, a multi-model ensemble review was tested. ReviewBench predicted it would achieve higher precision, recall, and comment volume with lower cost per review compared to the single-reviewer control.

These predictions held in subsequent online A/B tests. Recall rose by 13.6% and precision, measured as the addressed rate, increased by 8.0%. Comment volume grew by 61%. Also, the cost per review fell by 8.0%. Feedback in production shifted toward critical and moderate issues, with fewer minor suggestions. GitHub states that while ReviewBench helps identify promising updates, tests with users remain the final measure of impact.

Users can sign in to ReviewBench with GitHub and register an agent by providing its container image, configuration, and model access key to begin evaluation.

Related coverage

More from Vulnerabilities