TEST-TIME DILATION IN SEQUENCE MODELS
Abstract
A model can have access to earlier information yet fail to use it when a dependency becomes longer. We study test-time dilation: expanding the material between an earlier state and its use after training, while preserving the required answer. Controlled experiments separate three requirements: preserving one value, retrieving one of several bindings, and preserving an outer value while computing inside a scope. Attention can retrieve content directly, but positional scoring can make that route unreliable. Selective recurrence can preserve information, but repeated updates can alter it. Targeted interventions restore both forms of access, and scaling allows ordinary recurrence to learn more reliable carry. Neither successful retrieval nor accurate copying alone establishes robustness to the other requirements. A controlled scoped task makes the distinction concrete: models that master an order-sensitive program throughout the training range can still collapse when answer-preserving operations are inserted beyond that range. Dilation robustness is therefore a property of the learned computation, not a guarantee supplied by an attention edge or a long context window.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.