acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Optimization Is Not Test-Time Memorization: Auditing Neural Memory on Egocentric Video

Abstract

A test-time memory writes a stream into adaptive parameters by descending a self-supervised loss during the forward pass. This recipe can succeed as optimization while failing as memorization, and the training loss cannot tell the two apart. We audit a Titans-style MLP memory over frozen DINOv2 features. We first formalize when a recall objective cannot identify memory use, through a sufficient condition for a write-free global optimum and a Bayes-risk gap that states when causal history can help beyond the query. Trained for recall on i.i.d. image sequences, the memory approaches that optimum. The loss falls from 1.24 to 0.12 while the write rate contracts 200-fold, the final state loses input dependence, and retrieval drops below that of the same architecture at random initialization. Egocentric video adds a second shortcut. Frozen features remain correlated twenty seconds apart, and a causal convolution in the key exposes recent features to the query. Under a past-future windowed objective, the standard diagnostics look positive, and controls explain each of them away. At the default horizon, a copy of the oldest feature in the key window matches the far-past gain, a stateless filter with five coefficients per offset matches or beats the model at every offset, and the seen/unseen AUC ranges from 0.23 to 0.92 depending on which video is written. With a horizon beyond the key window, the model outperforms the filter twelve to seventeen seconds back while relying on its own state, yet its trace of individual items stays weak. The seen/unseen AUC detects neither case. We distill the audit into the controls that any storage claim about a test-time memory should report.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.