acceptodds
Under review as a conference paper at ICLR 2027

Rewards Beyond Test Passing: Towards Generalizable Fidelity in LLM Debugging

Abstract

Reinforcement learning (RL) from unit tests is the common recipe for post-training code models, but a unit test only asks whether the final program passes. We analyze what this unit-test-only reward teaches a model during debugging and find that, while it raises pass rates, it also teaches the model to edit more and to hallucinate bugs in code that is already correct. We introduce PreciseCoder, a training algorithm that augments the unit-test reward with a signal for semantically verified, localized edits, which correlates with both smaller edits and higher pass rates. On 4B and 9B models, PreciseCoder reduces the average number of edited lines from over 13 to 3 at a pass rate on par with unit-test-only training. Trained only on single-line bugs, PreciseCoder can generalize to repository-level program repair, where it resolves more issues with smaller patches. Its precision also transfers from weaker to stronger models: handing the reasoning trace of a 9B PreciseCoder model to frontier models such as GPT-5.6 Sol raises edit precision by 4.6 points and pass rate by 1.2 points with no additional training. Linguistic analysis of the thinking traces shows that our reward encourages the model to reason about individual lines, name the buggy ones, and form consistent bug hypotheses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.