Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families
Abstract
Benchmark scores describe current capability but do not determine how a checkpoint responds to further training. We introduce L-STATE, which combines current scores with capability changes from four short, target-independent training interventions. Its measured responses support direct and structure-preserving operator readouts, with a cross-family prediction bound that accounts for coordinate heterogeneity. Relative to current scores alone, both pulse readouts reduce prediction mean squared error by 39.4% across three development families. Reductions reach 78.3% for the operator on held-out GLM and 78.4% for the direct readout on held-out Granite. Fixed eight-step probes also predict longer responses: on nine GLM checkpoints under repeated-example adaptation, a post-hoc evaluation without refitting or rescaling retains 63.3% and 47.9% error reductions at 16 and 32 target steps. In a separate GLM/Yi comparison across repeated- and fresh-example training, full L-STATE reduces prediction error by 36.9–38.8% and action-selection regret by 24.8–68.8% at a common 32-step endpoint. These results show that short, reusable interventions provide information beyond static capability, improving both response prediction and training choices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.