Task identification features causally drive execution features in in-context learning
Abstract
Language models learn patterns from demonstrations in context, but how information at the demonstrations comes to drive the output has not been established. An identification feature for a pattern forms where the pattern is demonstrated (we call these “label tokens”) and drives an execution feature at subsequent locations where the model must continue the pattern (“cue tokens”). We establish two results: 1) an identification and an execution feature together mediate this effect on behavior, and 2) a single affine map shared across patterns relates each identification feature to its execution feature. In simple patterns generated by input-output tasks in GPT-J and Qwen2.5-7B, where the execution feature is the task's function vector, steering the identification feature raises cue-token alignment with the execution feature more than with a generic one, ablating it lowers that alignment, and ablating the execution feature removes the accuracy gained from identification steering. A single affine map fitted on training tasks predicts held-out tasks' execution features from their identification features ( of 0.60 and 0.53, versus below zero for a mean predictor). The same pathway appears in a naturalistic setting, in-context learning of coding conventions, with no explicit labels, features at evidence and cue tokens steer and ablate generated style, and identification steering moves the cue-token state toward the execution feature, although the linear map is weaker.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.