AJR: Amortized Jacobian Readouts for Prompt-Adaptive Steering
Abstract
Reading a language model's hidden state and predicting what will happen if we change it are different tasks. We introduce the Amortized Jacobian Readout (AJR), which learns from gradients to predict how a hidden-state edit changes the model's preference between two output tokens. Its prompt-dependent response direction both scores proposed edits and supplies a steering edit during the forward pass, without a backward pass at inference. Controlled color-retrieval experiments across four models from three families show why this distinction matters: learned decoders read unmodified states accurately but often mispredict intervention effects, whereas AJR learns their local response. In a fresh 640-prompt evaluation per model, steering with averaged linear and nonlinear AJR directions produces larger target-margin changes than fixed mean-gradient and contrastive-activation directions in every model. With the target answer supplied and edit size fixed from fitting data, it corrects 71–100% of initially wrong answers, while changing at most 2.7% of initially correct answers to wrong ones. Measured two-token generation latency is close to the unsteered model's; local-gradient steering takes 2.1–2.3 times as long. AJR thus turns learned intervention responses into prompt-adaptive control, within the depth and contrast on which it was fitted.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.