acceptodds
Under review as a conference paper at ICLR 2027

When Reasoning Contradicts the Table: Diagnosing Error Recovery in Language Models

Abstract

Language models often solve TableQA by generating reasoning traces that cite intermediate evidence. However, an incorrect table reference can make an otherwise coherent reasoning trace inconsistent with the evidence in the table. When such a conflict arises, can models return to the evidence and recover on their own? We investigate this with the *Evidence–Trace Conflict* (ETC) protocol: we hold the table, question, and continuation point fixed while varying whether the supplied reasoning prefix is correct or contains data referencing errors (DREs). Across model families and three TableQA benchmarks, conflicting prefixes cause accuracy drops of 49.4 to 78.7 percentage points. On WikiTableQuestions (WTQ), an explicit reconsideration cue placed after the conflict improves accuracy by 67.3 percentage points, from 18.0% to 85.3%, whereas a reminder in the original user prompt, before the conflicting prefix, provides little help. This contrast reveals a *self-initiation gap* under evidence–trace conflict: models often recover when reconsideration is triggered after the conflict, but rarely initiate it on their own. Guided by this diagnosis, we fine-tune models on self-generated ordinary and recovery continuations to internalize the externally elicited behavior. This training raises single-pass recovery from 18.0% to 88.0% on WTQ and achieves over 85% recovery on two unseen TableQA benchmarks. Together, these results show that when reasoning contradicts the table, models often follow erroneous traces despite having access to the correct evidence, but can learn to recover without additional inference-time intervention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.