From Smoothing to Noise Cancellation: Memory in Sequential Error Detection
Abstract
Language models often solve mathematical problems through a sequence of calculations and deductions. An evaluator, called a verifier, assigns each step a score intended to indicate whether the reasoning is correct. These scores are imperfect, so a detector must decide how to use earlier scores alongside the current one. Earlier scores may provide evidence, but they may also repeat the same scoring error. We study when these scores should be added as evidence, subtracted to cancel persistent noise, or ignored. We develop a theory of this choice for causal linear detectors under correlated noise, optimizing a class-weighted signal-to-noise ratio (SNR). In a stationary binary Markov model, all historical weights are positive when the state persists longer than the noise, negative when the noise persists longer, and zero when the two persistences match. After normalizing the current weight to one, nonzero total absolute historical weight strictly decreases as observation quality improves. We characterize when this law extends beyond binary states. For finite traces, initialization can reverse the stationary rule or produce mixed signs. A study of 27 methods on 2,000 MATH and GSM8K problems finds that the model-derived kernel improves within-position SNR over current-only scores in all six settings, with simultaneous intervals excluding zero. Matched-SNR comparisons and conditional-moment diagnostics assess when the model-prescribed weights are accurate. The results explain both the role of history and the conditions under which these memory rules apply.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.