ParetoMem: A Gradient-Descent Memory for Test-Time Streaming Context Encoding
Abstract
Memory mechanisms endow language models with learning from its own context stream, but test-time writing new context chuncks by gradient descent tends to erase what has been written before. We argue that this problem can be structural rather than only a matter of capacity: each write step optimizes only the loss of the chunk currently in view, while the evidence needed to defend earlier chunks has already passed, hence the two objectives are never weighed against each other. We therefore recast streaming writes as online multi‑objective optimization over a single shared state, in which every chunk already absorbed contributes an objective that the current update must not sacrifice. ParetoMem realizes this as a write rule instead of an architecture: it retains a small set of anchor gradients from earlier chunks together with a running summary of the rest, and projects each incoming update so that it improves the present chunk without measurably harming the retained ones, allowing each anchor a tolerance that grows with how stale it is. We prove that the rule is approximately Pareto‑safe, forgets a bounded amount per step, and cannot collapse the memory. It needs one fixed‑size buffer, independent of stream length, leaves the read path and parameters untouched, and can be attached to frozen pre‑trained backbones at test time — where it improves every metric on ten benchmarks at two scales, most on the benchmark most sensitive to losing early context, and by margins that grow as the stream deepens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.