acceptodds
Under review as a conference paper at ICLR 2027

DYNAMIC WRITES, SIMPLE SCHEDULES: AN ORACLE STUDY OF IN-PLACE TEST-TIME TRAINING

Abstract

Test-time training (TTT) adapts a language model as it processes an input, but existing methods typically apply a fixed adaptation configuration throughout the context. We examine whether varying where and when the model writes can outperform the best constant write action. Using a fixed Qwen3-1.7B In-Place TTT checkpoint, we compare the per-example best constant action against a bounded oracle search over chunkwise dynamic trajectories. Across 6,500 examples from 13 RULER-16K tasks, the best trajectories returned by the bounded oracle search increase the mean score from \(78.38\) to \(79.77\), a gain of \(1.39\) percentage points. Among the 1,933 non-ceiling examples, the bounded oracle search improves 170 and ties the remaining 1,763, increasing their mean score from \(27.29\) to \(31.96\) (\(+4.67\) points). Controlled interventions show that successful outcomes depend on both the ordering of write actions and the model state accumulated across chunks. Nevertheless, complex switching is often unnecessary: 62 of 78 improved multi-switch trajectories admit an exactly score-matched one-switch alternative. We observe similar, though smaller, gains on Qwen3-4B and Llama-3.1-8B under their respective 32K-token protocols, improving over the per-example best constant action by \(0.79\) and \(0.36\) points, respectively. These results establish an oracle-reachable advantage for dynamic in-place TTT and suggest that relatively simple switching policies may capture much of this opportunity. It motivates learning policies that select write actions from the evolving context and adaptation state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.