acceptodds
Under review as a conference paper at ICLR 2027

Width or Interaction Fidelity? Allocating a Fixed State Budget in Analytic Continual Learning

Abstract

Analytic continual learners fit a closed-form ridge readout on frozen features and carry only its sufficient statistics, the continuation state, across tasks. Expanding features into many random coordinates improves accuracy, but the Gram matrix in this state grows quadratically with width. We ask how a fixed continuation-state budget should be split between representation width and the fidelity of stored feature interactions. Because the feature-label statistic and the Gram diagonal grow only linearly with width, we keep them exact and compress only the interactions with **Diagonal-Preserving Spectral Ridge (DPSR)**, which merges per-task spectra into a rank-bounded factor and solves a full-width readout in closed form. On three class-incremental benchmarks with ViT-B/16 features, DPSR stays within points of a full-Gram readout of the same expansion in stage-average accuracy while storing less continuation state. At equal budgets of MiB with penalties chosen by generalized cross-validation (GCV), trading interaction rank for width beats exact ridge on a narrower expansion in of configurations by points (three seeds, paired intervals above zero) and loses points where too little rank remains. With validation-and-refit selection, gains are smaller, clearest on CIFAR-100 and shared by other corrected sketches, and a penalty chosen before truncation can collapse accuracy after refitting. Under GCV, width thus outweighs interaction fidelity down to a dataset-dependent rank; the savings concern persistent state, not training time or peak memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.