acceptodds
Under review as a conference paper at ICLR 2027

Consistent but Wrong: Analyzing What LLM Reasoning Lacks via State Decoding

Abstract

Language models are increasingly used for multi-step reasoning. They decide at the token level, and every later step builds on what they have written. Sampling-and-selection (best-of-N, majority voting) chooses among chains already written. A *consistent error* occurs when all sampled chains that reach a decisive step write the same wrong successor state, so the pool holds no correct candidate at that step. We use *state decoding* to analyze what the model's reasoning lacks at such a step. Its *support* supplies the candidate states at each step, the *scoring* ranks them, and the *search* assembles the decisions. An enumerated support contains the correct successor whatever probability the model assigns to it, and with an exact search every remaining error is a scoring error. With Qwen3-4B, we run SAVI (State-Aware Viterbi Inference), which enumerates the candidates and searches exactly on *MuSR* (belief tracking) and *Belief-R* (belief revision). In both tasks, most failures that more samples do not repair are consistent errors. Once the candidates contain the correct state, the model's likelihood picks a belief table on MuSR by the sentence that asserts the move, not by who saw it. On Belief-R it picks a reading by its wording. When the prompt states the correct reading, the model answers correctly, so it fails at choosing the reading. For Qwen3-4B, only signals outside the likelihood pick the correct candidate above chance. Of 19 MuSR items on which voting stays wrong, a verifier built from the annotated facts fixes 15 when it scores enumerated states and 3 when it scores the sampled pool. On *Countdown* the chains disagree, and frequency does not track target reachability: of 45 items on which every sampled chain fails, state decoding solves 9 with frequency scoring and 44 with an exact verifier. With uniform proposals and the same verifier, it solves 39 to 42. Separating the candidates without labels or an exact verifier remains open.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.