acceptodds
Under review as a conference paper at ICLR 2027

Better Risk Hypotheses, Better Code Reviews? Diagnosing Risk Formulation and Prioritization in LLMs

Abstract

AI-assisted coding can accelerate code production, increasing the need for efficient and reliable code review. Yet evaluating LLM reviewers through final findings alone obscures whether failures arise from overlooking relevant risks, prioritizing unproductive investigations, or failing to verify a concern. We introduce a diagnostic framework centered on Risk Hypotheses: structured, testable descriptions of potential failures and the evidence needed to investigate them. From real review histories, we curate 300 reference risks grounded in reconstructed review-time evidence, distinguishing concerns worth investigating from those verifiable in a fixed environment. The framework evaluates risk formulation, budgeted prioritization, and risk-guided review. Controlled comparisons isolate selection effects using fixed candidate pools and verifiers, while Gold-derived guidance and field ablations probe how hypothesis content affects downstream findings. Initial experiments show that broader risk coverage does not necessarily improve top-ranked coverage, while Gold-derived guidance can improve finding recall for some reviewers. We also identify grounded risks that require evidence unavailable in the review environment. Together, these observations show why recognizing a relevant risk, prioritizing its investigation, and substantiating a finding require separate assessment. Our framework connects these intermediate capabilities to review outcomes, enabling systematic diagnosis of when better risk hypotheses lead to better code reviews.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.