The Reach of Interventions: What Can Be Written into a (Semi-)Frozen Language Model?
Abstract
Writing new knowledge into a partially frozen language model means changing an answer's reachability through one of three routes: carrying and placing content into weights, carrying and placing content through activations, or reweighting or tilting the base model's existing sampling support. Fine tuning, RL, steering, retrieval, and editing span these routes, and the route determines what each can reach. We develop a predictive theory of what each intervention can write. Using sampling support, verifiability, and a prior-support decomposition, the theory separates items into display, verifiable closure, and underdetermined layers, and derives an inheritance law for reachable answers. It predicts learning limits and per item learnability before training. Across synthetic and public benchmarks, the theory explains channel specific reach and failure modes, matches observed ceilings, and turns post training from trial and error into measurement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.