acceptodds
Under review as a conference paper at ICLR 2027

Helpful When Right, Controlling When Wrong: Causal Dependence on External Intermediate States in Small Language Models

Abstract

Step-level feedback is usually judged by whether it improves a final answer, but this hides whether a model recomputes or simply inherits a supplied intermediate state. We isolate these mechanisms with controlled interventions on arithmetic and symbolic reasoning traces. Each trace has a unique first error and deterministic continuation, so we can change only the external state while measuring final repair, local adoption, and exact downstream propagation. Across Qwen3 and Gemma 3 models from 270M to 4B parameters, correct states sometimes produce large gains. On symbolic logic, they improve repair over localization alone by 58.0 percentage points for Qwen3-1.7B and 18.7 points for Gemma 3 1B. Capable models also propagate false states. Qwen3-4B adopts matched wrong feedback on 51.3% of arithmetic items and 56.7% of logic items. A pre-specified seven-level arithmetic intervention reveals a sharp peak at the true state, alongside persistent family-specific susceptibility to counterfactual states. The results support three descriptive regimes: capacity limited, feedback enabled, and autonomous but feedback susceptible. External step-level feedback can therefore be both helpful and behaviorally controlling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.