Training Selects Among Behaviorally Equivalent Computations in In-Context Dictionary Learning
Abstract
Behavioral agreement with a known algorithm does not show which computation a network implements. We make this question exact in in-context complete dictionary learning, where a global inverse, a support-masked inverse, and a support-gated local decoder are provably identical on every training query and differ only beyond the sparse set. A sparse-trained Transformer does not learn the global inverse: fed queries denser than its training set it returns sparse-shaped outputs, and its derivative on held-out directions lies nearest the gated-local decoder, the minimum-norm endpoint of the equivalent class. Training design, not the task, sets where in that class the network lands. At matched sparse-task accuracy, removing the auxiliary support losses of the standard objective moves the learned map toward the support-masked inverse and widening the training query support moves it toward the global one. One support step or a 0.5% dense-query admixture removes most of the sparse-to-dense performance gap; the admixture acts within hundreds of steps on a trained network and reverts once withdrawn, while the derivative on sparse queries keeps its position in the class. As dictionaries grow, errors concentrate in selecting the support rather than in decoding on it. When supervised behavior leaves the computation unidentified, training design selects the extension.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.