acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Revise After a Prediction Error

Abstract

We investigate sequential diagnosis, considering whether models can use new evidence to distinguish a faulty observation from a changed state or governing rule, and whether lightweight fine-tuning on synthetic examples can improve their ability to identify what revisions benefit real scenarios. We build matched triplets in which the three causes are indistinguishable until later evidence arrives, with an exact Bayesian reference at every step, and find that language models compute the diagnosis but do not report it. LocusProbe, a linear probe on Llama-3.1-8B's hidden states, identifies the cause on real weather series with 0.97 accuracy against 0.44 for the same model's answers, and beats every tested open direct-answer system on matched items, including the nine-times-larger Qwen2.5-72B by 0.125 [0.080, 0.173]. Posterior-matched tuning (PMT) on synthetic arithmetic and logic alone transfers to real weather and industrial series with no real labels; applied to the unembedding only, with every other weight frozen, it raises rule-change recall from 0.03 to 0.72. When the diagnosis selects the revision, probe-selected revisions recover 98.5–99.9% of the gap to an oracle over 60 steps and attain a worst-case loss of 6.63 against 14.98 for an observer given the exact generative family of all three causes. Recall on rule changes, not overall accuracy, predicts that recovery ( versus ), because an uncorrected rule compounds. These results identify a learnable component of revision, i.e., using subsequent evidence to determine what went wrong, what should change, and whether that revision allows the system to recover.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.