How Test-Time Updates Become Behavior: Objectives, Training Scope, and Update Paths
Abstract
Can large language models turn information in deployment-time support into behavior that remains accessible after the support is removed? We study a Support-Derived Behavioral Objective (SDBO) that constructs one question–answer target from the support and a fixed schema, then takes four local stochastic-gradient steps. On 32 independent QAEdit confirmation episodes, SDBO improves joint attained success for Llama-3.1-8B-Instruct over raw-support training by 37.5 percentage points at both primary KL budgets. Paired analyses establish an objective-by-training-scope interaction; content controls and held-out formulations show that the effect depends on the correct support content and extends beyond the training wording. Two fixed update policies achieve positive native-query net gains and pass relative-preservation non-inferiority tests. To probe why scope matters, we decompose full-model updates and intervene on 32 separate fresh episodes. Scaling each full-update projection to the effective radius of its independently trained local counterpart leaves a 62.5-percentage-point behavioral gap at both policy points. Path comparisons further support a learning-rate-driven multistep contribution. The complete gain-and-preservation pattern is not reproduced on Qwen3-8B-Base, marking a model-setting boundary. These results show that final parameter support and update magnitude do not determine behavior: the process that forms the update also matters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.