Why Errors Survive Selective Grounded Correction
Abstract
Large language models (LLMs) can make mistakes even when given the sources needed to answer. A support score is the fraction of output claims fully supported by those sources. But does a higher score mean more mistakes were fixed? Across legal, document, and open-domain question answering, we study selective correction, which edits only claims a checker marks as errors. Holding the first edit fixed, we compare removing claims that still fail with editing them again. We find that an answer can say less and score higher, or fix a mistake and keep the same score. Editing again restores support to 37 claims that removal would discard, yet changes average support by only to percentage points. Weaker wording can also raise scores: adding "It appears that" before checking and removal raises support by 2.0–2.5 points over removal alone in our legal experiments. Human annotators find that all 52 sampled hedges narrow the assertion and that some credited repairs edit already supported claims. In 17 of 18 traced configurations, most surviving errors are not sent back for another edit, and failed repairs dominate the other. We call for reporting what each step repairs, preserves, and removes, with its token cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.