AdaWrite: Adaptive Write Depth Routing for Test-Time Training
Abstract
Test-time training (TTT) models retain information from long contexts in fast weights that are updated during inference. Common chunkwise implementations process each chunk at full depth, even though the benefit of deeper processing can vary. We introduce AdaWrite, a parameter-efficient method that reduces inference computation during context processing, or prefill. AdaWrite learns where each updating chunk stops in a frozen TTT backbone. An early exit skips deeper computation and memory updates while leaving skipped layers' fast states unchanged. We first train intermediate readouts, then learn the routing policy with a detached full-depth teacher as a reference and a compute penalty. Our main experiments on LaCT cover retrieval, document understanding, and language modeling. On S-NIAH-1, AdaWrite substantially reduces analytical prefill FLOPs and improves retrieval beyond the training context length. The router concentrates deeper updates near the needle without needle supervision. On other retrieval tasks and natural text, the same policy uses more layers and achieves prediction quality close to the full-depth baseline. An additional study on TTT-E2E shows modest computational savings with a different memory update rule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.