GitHub launches ReviewBench open benchmark to evaluate AI code reviewers
GitHub has launched ReviewBench, an open benchmark created to test whether AI code reviewers catch meaningful issues without adding noise. The benchmark is derived from an analysis of 103.9 million GitHub pull requests and compiles 219 pull requests across 19 programming languages.
Developers can bring their own code review agents, evaluate them against the benchmark, and submit results, as detailed in [GitHub's blog post](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/?utm_source=x-promoting-reviewbench-blog-article&utm_medium=social&utm_campaign=reviewbenchmark-oct-2026).