acceptodds
Under review as a conference paper at ICLR 2027

The Reach of Interventions: What Can Be Written into a (Semi-)Frozen Language Model?

Abstract

Writing new knowledge into a partially frozen language model means changing an answer's reachability through one of three routes: carrying and placing content into weights, carrying and placing content through activations, or reweighting or tilting the base model's existing sampling support. Fine tuning, RL, steering, retrieval, and editing span these routes, and the route determines what each can reach. We develop a predictive theory of what each intervention can write. Using sampling support, verifiability, and a prior-support decomposition, the theory separates items into display, verifiable closure, and underdetermined layers, and derives an inheritance law for reachable answers. It predicts learning limits and per item learnability before training. Across synthetic and public benchmarks, the theory explains channel specific reach and failure modes, matches observed ceilings, and turns post training from trial and error into measurement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.