RESEARCH · RESEARCH · #1547
ReviewBench: an open benchmark for AI code review
ReviewBench is a new publicly available offline benchmark for AI code review agents that contains 219 public pull requests across 19 languages, modeled on GitHub distributions and validated by senior engineers. It builds a multi-source golden set (human reviewers, follow-up commits, static tools, and multiple LLMs), uses Claude Sonnet 5 as a consistent grader, and provides metrics and labeled findings (severity and category) to compare reviewers and track likely production impact.
KEY POINTS
- ReviewBench is a new publicly available offline benchmark for AI code review agents that contains 219 public pull requests across 19 languages, modeled on GitHub distributions and validated by senior engineers.
- It builds a multi-source golden set (human reviewers, follow-up commits, static tools, and multiple LLMs), uses Claude Sonnet 5 as a consistent grader, and provides metrics and labeled findings (severity and category) to compare reviewers and track likely production impact.
- A realistic, validated benchmark with diverse ground truth and production alignment helps teams compare AI code reviewers meaningfully and predict which improvements will matter in real workflows.
WHY IT MATTERS
A realistic, validated benchmark with diverse ground truth and production alignment helps teams compare AI code reviewers meaningfully and predict which improvements will matter in real workflows.