VeriLoop Structural Repair: Quantifying Progress Without Regression in Code Repair
Abstract
Test-based program repair can hide damage to previously correct behavior when it reports only aggregate test success. We introduce VSR, which constructs training and evaluation data from protected-obligation rank vectors and the non-regression relation . The dataset contains 40 program families, 988 tasks, 2,297 states, and 27,046 executed candidates and passes 228,473 independent checks. Base/SFT comparisons at both 4B and 9B show that supervised repair data improves closed-loop completion. Construction ablations show that coupling-aware sampling exceeds equal-size random sampling at both model scales and on all three test distributions, while the full SFT mixture similarly exceeds COMPLETE-only in every matched comparison. Sequential 9B SFT-to-DPO raises completion by , , and on ID, OOD-family, and OOD-, respectively. In the DPO error-type ablations, removing either aggregate-trap or protected-regression examples causes a comparatively large drop in completion and usually lowers NRIR or increases calls, highlighting their roles in resisting aggregate-score shortcuts and preserving previously correct behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.