Content-Dependent Recovery from Task-Relevant Activation Disruption in Large Language Models
Abstract
When prompts are modified or rewritten, large language models (LLMs) need to determine which information in the prompt remains valid in order to continue completing the specified task. Prior work has examined how models process such changed information, but what happens to the model when its internal computation is disrupted remains largely a black box. We gradually attenuate internal activations at task-related token positions, select intervention strengths from the resulting performance curves, and write activations recorded during normal runs back into the perturbed model. Experimental results across five open-weight models and four reasoning tasks show that reasoning performance drops more when task-related information is disrupted, while nearby layers can show similar effects. When activations carrying the correct information are written back, the models can recover their reasoning performance. We further find that recovery in reasoning does not simply correspond to internal activations returning toward their normal state: incorrect information can also move perturbed activations closer to the normal state while producing substantially weaker recovery of the receiver's correct answer. This pattern is observed across all five models, while component analyses show model-dependent effects, with the clearest effects occurring at different attention and feed-forward sites.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.