Prediction Is Not a Sufficient Memory State
Abstract
A memory that reproduces every current prediction may still have forgotten how to learn from the next observation. We characterize this distinction for exact cumulative learning and attention. For smooth convex residual losses with unrestricted real features and labels, a continuous recurrent state of fixed dimension exists exactly when the loss is polynomial. With ordinary, unweighted observations, degree requires precisely coordinates; even among histories with one fixed current predictor, exactly coordinates can remain necessary for continuation. Nonpolynomial losses have unbounded requirements on such equal-prediction histories. Bounded scalar logistic examples retain this obstruction after fixing any finite derivative jet. For ridge regression, the hidden evidence is the design matrix: the exact state costs , versus for the predictor. On a well-conditioned zero-prediction family, optimal -bit compression has average one-update error of order , , even with arbitrary randomized codes. Kernel attention gives a second exact closure classification. The results identify what future updates reveal, rather than proposing a new regression algorithm or asserting trained-model gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.