Computation Before Correction: Separating Diagnostic and Solution Value in LLM Reasoning
Abstract
Test-time reasoning is shifting from simply scaling computation to deciding when and where additional computation is worth spending. We ask whether reasoning can be useful before it can solve the problem: can short computation reveal whether substantially more, independently restarted reasoning is likely to succeed? We study this as diagnostic computation using a probe-and-discard design that separates solution value from observer-relative diagnostic value. From a frozen wrong reasoning state, a short probe is generated, observed, and discarded; evaluated long repairs restart independently from the original pre-probe state, preventing probe tokens from contributing to the repair itself. In a prospectively frozen Qwen3-8B/MATH experiment, a 16-token probe directly corrects only 1.50% of cases yet substantially improves prediction of future repairability. A separately frozen DeepSeek/GSM8K experiment reproduces the qualitative separation across model family and benchmark. More broadly, the results suggest that test-time computation can play two roles: advancing a solution or measuring whether further computation is worth spending. This motivates treating solution value, diagnostic value, and decision value as distinct quantities rather than assuming that useful computation must immediately improve the answer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.