acceptodds
Under review as a conference paper at ICLR 2027

BEYOND UPDATE COUNTS: A CONTROL MAP FOR EVALUATING TEST-TIME MODEL UPDATES

Abstract

Test-time updating lets a language model change persistent state after training—such as model weights, adapter parameters, or a memory module—before or while answering queries. Existing evaluations often compare systems using update count or frequency and performance measured immediately afterward. These summaries omit what is updated, where, when, and in what order. The same total can therefore hide an important difference: one system may answer after one update while another answers after two. We introduce the Control Map, an evaluation framework that records the committed updates preceding each scored answer and separately evaluates behavioral changes, immediate learning (acquisition), later recall (retention), preservation of unrelated behavior (locality), and performance in new settings. We prove when counts alone determine final state and show, under stated assumptions, how differences between pre-answer histories limit differences in average scores. Our experiments establish three lessons. First, choices hidden by update count matter. At one 760M-parameter checkpoint, 24 equal-count schedules span 19.7%–34.2% accuracy; updating the early and late layer groups more often than the middle groups performs best in this experiment. Second, apparent count effects can disappear after proper matching: in 1,600 pairs of runs assigned different nominal dose labels but executing the same updates in the same order, both runs end with the same model state and locality. Third, immediate learning does not guarantee a lasting or well-contained change. Across six selected CounterFact LoRA settings, new information is acquired in all six; in a separate ROME study, edits are retained at all 13 tested operating points but fail locality at all 13. The relative ordering of editors also changes across the tested combinations of models, benchmarks, and tests. Therefore, evaluations should record which updates precede each answer and assess these questions separately.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.