What You Don't Check, Models Don't Fix: False Completion under Partial Verification
Abstract
Iterative correction loops revise LLM outputs until their checks pass, and self-improving pipelines favor changes that score better on theirs. However, the checks rarely cover every way an output can be wrong, so an output accepted as correct may well contain defects that the checks missed. We ask what happens to these defects. We study a correction loop on three task domains (music-as-code audio, SVG repair, and API workflows) with eight proprietary and open-weight models. Each task has eight deterministic checks. All eight checks are scored in every round, but the model sees only some of them (the exposed checks; the rest are hidden), and the loop stops when the exposed checks pass. Because the hidden checks are also scored, each run shows which defects the loop repairs and which it leaves behind when it stops. First, models fix what they are shown: exposed defects are repaired far more often than hidden ones, and at low coverage most of the loop's stops leave a hidden defect, usually before any revision. Second, stopping too early is not the main reason these defects remain: extra rounds recover little, and telling the model that its output is still wrong helps, but far less than naming the failing checks. Third, models can compensate only partly: reasoning reduces the effect of coverage without removing it, and letting models choose which checks they see does not detectably reduce the remaining defects. What you don't check, models don't fix: a caution for any correction loop trusting partial verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.