Diagnose, Then Select: What Makes Critique-Based Refinement Effective?
Abstract
Large language models increasingly rely on critique to refine their reasoning, yet existing approaches treat critique as an undifferentiated block of feedback and provide more information than may be necessary. We argue that critique-based refinement has two separate bottlenecks: the feedback must carry the causal information needed to repair an error, and the resulting revision must be reliably accepted or rejected. First, we investigate which parts of a critique make an error repairable by decomposing one structured record into isolated correctness-verdict, error-location, and error-diagnosis views, and compare them with the full critique. Across three refinement protocols, two model families, and two mathematical-reasoning benchmarks, diagnosis-only feedback recovers of the wrong-to-correct rate of a full critique while using of its feedback tokens; verdicts and locations recover only and . Correction alone, however, is an incomplete objective: critique also destabilizes answers that were already correct. Under solver–critic refinement on MATH-500, diagnosis corrects of the weaker Llama solver's wrong answers but corrupts of its correct ones, a net loss of points, while the same feedback gains points for the stronger Qwen solver. To address the second bottleneck, we propose **DiSel**-Consensus, a sampling-based revision selector that accepts a diagnosis-conditioned revision only when one independently sampled auxiliary revision supports it. At a common calibration risk budget, Consensus reduces correct-to-wrong transitions from under unconditional revision to , retains wrong-to-correct transitions, and gains accuracy points with 923 marginal tokens per trial. Compared with **DiSel**-Pointwise selection baseline, **DiSel**-Consensus achieves a significantly lower corruption rate, while the difference in net accuracy is not statistically significant.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.