acceptodds
Under review as a conference paper at ICLR 2027

Beyond Recoverability: Understanding Failures in Intermediate Reasoning

Abstract

Intermediate representations are widely used to support reasoning in large language and multimodal models, yet they are typically evaluated by final-answer accuracy, which provides little insight into why a representation succeeds or fails. We introduce PADC, a diagnostic framework for localizing failures along the representation-to-answer pathway by separating Preservation, Access, Downstream Dependence, and Correct Computation. The framework is designed to distinguish whether failure arises because task-relevant information is not retained in the representation, cannot be reliably recovered by the model, fails to control downstream prediction, or is available and influential but still not transformed according to the task rule. Across controlled synthetic tasks and six natural-domain benchmarks, matched interventions reveal distinct signatures at these interfaces. Removing task-relevant structure produces strongly heterogeneous effects across tasks, indicating that relevance alone does not determine downstream necessity. In controlled relational reasoning, local task state can be recovered perfectly while composition over that state remains unreliable, separating state access from successful downstream integration. Most notably, held-out donor-state interchange induces systematic prediction shifts beyond matched sham interventions, yet the model rarely remains correct under both original and counterfactual states, showing that causal sensitivity to a recovered state is not sufficient for task-correct computation. The same qualitative separation persists across additional model families. Our results suggest that intermediate-representation quality should not be reduced to final accuracy, information retention, or recoverability alone; a useful reasoning representation must expose task-relevant state in a form that not only affects prediction, but supports the computation required to produce the correct answer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.