Test-Time Generation
Abstract
Long-context models increasingly rely on fast-weight and test-time-training memories—state-space and linear-attention layers that compress a growing context into a fixed-size state and recall it in linear time. But they read that state through deterministic regression, which under squared loss maps a query key to the conditional mean of its bindings and averages competing values away. Our contribution is to recast the readout, not the memory, as conditional generation: casting fast weights as a conditional associative memory, we show that a key-conditioned denoising readout augments key-space similarity with a value-space posterior term—the term a squared-error endpoint map discards. We realize this as Test-Time Generation ( TTG ), a drop-in readout that leaves the write operator and fixed-size state untouched and so stays orthogonal to the memory architecture: online updates supervise noisy versions of the observed bindings, and a query is answered by sampling a short denoising trajectory rather than regressing to a point. A compressed memory then need not recall by averaging to a single endpoint, but can use the cue to sample the binding it stored. A controlled two-values-per-key probe isolates the mechanism: no deterministic readout—not even one granted a much larger memory—recovers both competing values on any key, whereas the generative readout does at low retrieval load, and freezing its sampling noise removes the effect. The same readout improves both long-context language modeling and novel view synthesis over deterministic baselines, without enlarging the memory or its peak footprint, at a modest readout cost. On language modeling a single readout step already suffices for the gain, so readout depth is matched to training rather than scaled at test time. TTG thus turns a fixed limitation of test-time memory into a readout we can design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.