acceptodds
Under review as a conference paper at ICLR 2027

How Instruct Models Inherit Reasoning

Abstract

Instruct models increasingly read prompts written by other agents, as systems delegate simpler tasks to cheaper models. These instructions are seldom verified, but they are often the backbone of complex agentic operations. Nobody needs to verify them, as long as two assumptions are true: a model good enough at the task will not be misled by errors in the input, and if it is misled, its output will show it. We find that neither holds. Using BIG-Bench-Hard tasks rebuilt as generators, across twelve difficulty tiers and five models, we show that instruct models treat worked solutions as the answer, and therefore errors in the solution are propagated in the large majority of generations; for the hardest tiers the same holds when the error is planted mid-derivation, where no answer is stated anywhere. The rate at which errors are propagated does not fall as the model gets better at the task, and it leaves no trace in the answer. We study the copying mechanism using activation patching, and show the provided value is transported to the position that emits the answer. Hence, errors in the computation an agent passes on are not repairable by a stronger model or by reading its answers. We also study strategy prompts, in which an agent gives a task-specific overview of how the task is to be solved. A correct strategy buys about 40% of what a full solution buys, at a fraction of its cost. A strategy for the wrong task is ignored by some models and derails others, and a plausible but wrong strategy is followed at rates set by the model rather than by capability. Switching the same weights to thinking mode removes much of the copying, but less so on the hardest problems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.