acceptodds
Under review as a conference paper at ICLR 2027

ClaimSift: Selective Answer Revision in Large Language Models under Mixed Feedback

Abstract

Large language models (LLMs) can revise answers using feedback. With mixed correct and incorrect feedback proposals, however, successful revision requires judging the proposals, updating dependent items, and preserving initially correct unrelated items. Final Accuracy measures the revised answer but obscures these decisions. Moreover, complete revisions with the same revision reward can cause different damage to initially correct items. This raises the question of how to learn selective revision while accounting for this damage. We introduce ClaimSift, a ProofWriter- and FOLIO-based benchmark for evaluating complete revisions, and propose Preservation-Aware Revision Optimization (PARO). PARO combines revision reward with explicit preservation cost, while a shared penalty weight adjusts to observed risk relative to fixed risk budgets. Experiments on ProofWriter show that PARO achieves the highest Overall Success among evaluated Qwen3-8B trained revision methods, exceeding TD-DPO and DPO by 9.75 and 9.59 percentage points, respectively, and improving Content Preservation over DPO by 3.42 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.