Residual Re-entry Dependence: When Suspended States Are Not Sufficient for Resumption
Abstract
Task interruptions are increasingly common in language-model reasoning, but it remains unclear what information is sufficient for a model to resume a suspended computation. A natural assumption is that once the suspended task state and the next operation are known, the continuation is determined. We test this assumption and find that it can fail: even when the suspended state and continuation operation are held fixed, changing the intervening computation can systematically change the resumed result. We call this remaining influence residual re-entry dependence. To measure it, we construct controlled task-switching experiments that hold the suspended state and continuation operation fixed while varying the intervening computation. In a controlled arithmetic task, inserting an intervening computation reduces state-report accuracy from 100% to 76.67% and continuation accuracy from 100% to 72.85%. The two failures are not the same: in 17.99% of interrupted trials the model reports the suspended state correctly but continues incorrectly, while in 14.17% it reports the state incorrectly but nevertheless continues correctly. In a strict within-example factorial, simply changing the relation between the intervening and resumed operators changes continuation accuracy: exact operator matching yields a +6.08 percentage-point contrast, with the same direction on Llama-3.1-8B (+4.95 pp) but the opposite direction on Qwen2.5-3B (−10.42 pp). Thus, knowing the suspended state alone does not explain how the model resumes. Matched controls rule out two simple explanations: making the intervening numerical state conflict with the suspended state does not consistently worsen continuation, and explicit task-closure cues perform no better than matched filler text. We also find that interruptions broadly change how the resumed computation attends to the suspended-state span, but late-layer activation-patching effects that appear to transfer the state are largely explained by answer identity under same-answer controls. The same pattern appears in a natural reasoning task. On GSM8K, inserting an unrelated factual question changes final-answer accuracy little, whereas inserting another math problem reduces continuation accuracy by 3.2–6.1 percentage points in five of six models; the corresponding state probe degrades in all six. By contrast, in a symbolic permutation task, an apparent interruption effect disappears once operation priming is matched. Together, these results show that the suspended task state is not always a sufficient statistic for resumption: even when the state and next operation are fixed, what the model computes during the interruption can still change what happens after resumption.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.