Minimal Functional Modification: Minimal Behavioral Realization of Activation Interventions
Abstract
Activation interventions provide one internal realization of a desired language-model behavior, but need not provide a minimal one. We study minimal behavioral realization: given a frozen reference intervention function and a specified behavioral observation, what is the smallest hidden-state modification required to reproduce the reference behavior across prompting contexts? This problem is context dependent because demonstrations alter both model states and local sensitivities to intervention. We first characterize when directional modification is strictly more economical than optimal scalar contraction. More importantly, we show that pointwise redundancy need not be shareable across contexts: the excess cost of a shared finite-width realization is governed by the spectral structure and input predictability of the pointwise optimal intervention field. We further quantify the cost of preserving richer behavioral observations and establish conditions under which local advantages persist at finite intervention amplitudes. These results motivate Minimal Functional Modification (MFM), a shared nonlinear realization module trained directly to balance reference-output fidelity and hidden-state displacement without explicitly constructing Jacobians. Across two language-model families and three multiple-choice reasoning benchmarks, MFM reduces reference intervention displacement cost by 57.6% on average while preserving four-shot answer distributions. Across changed four-shot demonstration sets, MFM yields contextual-response errors of 0.0085–0.0103 across the six model–task settings. When answer-set probability mass is additionally constrained, MFM still achieves 43.4% displacement-cost savings on the evaluated subset.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.