Lego-Review: Scaling Code Review Data with Verification from Pull Request Histories
Abstract
AI coding agents produce a growing share of software changes, making automated code review increasingly important. A promising way to train capable reviewers is to scale the tasks they practice on and the verified trajectories used for training. However, code review yields open-ended, natural-language findings, for which there is no general test-based verifier comparable to the tests used for feature implementation or issue resolution. We observe that pull request (PR) histories offer an implicit verification signal: discussions, code revisions, and other actions can show which findings were accepted or fixed and which were withdrawn or rejected. We introduce Lego-Review, a pipeline to turn this implicit evidence into verifiable review tasks and filtered trajectories. We collect merged GitHub PRs, reconstruct the code snapshots reviewers examined, and use a curation agent to consolidate findings across each PR's discussions and revisions. The agent labels each finding as supported, opposed, or uncertain and binds it to the reviewed snapshot. A finding-level verifier then filters teacher trajectories by crediting recovery of supported high-priority findings and penalizing repetition of opposed ones. From 357k PRs, we build 112k review tasks in eight languages and 24k filtered trajectories. On our proposed Lego-Review-Bench F1 rises from 5.7 for the base model to 43.1 with 1k filtered trajectories and to 52.9 with 16k, showing effective scaling with verification from PR histories. We release the data, the benchmark, and the trained reviewer to support open research on scalable code review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.