acceptodds
Under review as a conference paper at ICLR 2027

Parameter Patching: Causal Analysis of Learned Computations in Language Models

Abstract

Large language models increasingly solve complex reasoning and coding tasks, yet learned computations supporting these capabilities remain poorly understood. Activation-based methods, such as activation patching (AP), reveal how internal states contribute to behavior; however, they leave open how these capabilities arise from the model parameters themselves. To study this dependence directly, we introduce **Parameter Patching (PP)**, which probes components through fixed edits to their parameters. As the parameter counterpart of AP, PP modifies the learned transformation applied throughout execution, rather than substituting the activation states it produces. Across tasks of varying complexity, we find that both methods locate components with causal influence on behavior, yet the components they select serve different behavioral functions. On classical interpretability tasks such as IOI and arithmetic, the two methods indeed select heads with distinct characteristics. For instance, on IOI, PP selects heads that more strongly affect the preference between candidate names than those selected by AP. This divergence matters more on complex reasoning and coding tasks, where PP selections generally incur larger accuracy losses under matched knockouts, indicating that they identify components on which these capabilities more strongly depend. Finally, PP selections also translate into practical gains. Fine-tuning PP-selected heads yields larger improvements than fine-tuning AP-selected ones on both mathematical reasoning and agentic coding tasks, and the same heads support effective activation steering. Together, these results show that intervening on learned transformations reveals behavioral dependencies complementary to activation-based analysis and identifies useful targets for improving model capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.