acceptodds
Under review as a conference paper at ICLR 2027

Depth is Degree: Test-Time-Training Updates Are Polynomial Filters, and Learned Step Sizes Only Choose the Coefficients

Abstract

We replace the learned constants of test-time-training updates — step sizes in TTT, momenta in Titans, conjugate-gradient coefficients in MesaNet, the single step of linear attention — and of unrolled restoration solvers with a polynomial head that computes them from the operator: it starts at the textbook solver, costs the same operator applications, and learns no solver constant. The head rests on one fact: each such update is a solver for a quadratic in-context loss, and we prove it is a polynomial filter in its operator whose degree is the number of operator applications, so learned step sizes and momenta choose a filter but cannot enlarge the class; the head computes the class-optimal one. The regime is one protocol on one GPU (42M and 125M parameters), where the ordering does not change with size. On multi-query associative recall the head ties learned step sizes, momentum, and conjugate gradient at the ten-seed median (86/57 at loads 32/64); the tie is the claim, and the head is the only rule with no failed seed. Inside language models it is the lowest-perplexity update at degree four, by 1.8 over learned step sizes at 125M; on in-context regression with a dominant nuisance mode — a high-pass target — the exact solve and adaptive conjugate gradient fail at 2.7× the head's error, which is the class optimum. The theorem is also a design rule: capacity enters only through the nonlinearities and re-estimated operators between solver steps — attention is one — and the solver constants choose within a class the head spans directly, so they can be computed rather than learned. In one unrolled solver network whose operator is re-estimated in every layer (image restoration) the bounded head ties the textbook solver with no rate to choose and no failed seed, while learned step sizes collapse at three of four solver rates and at the fourth land above the exact solve of their own system — a gain the theorem assigns to the re-estimated operator, not to the constants; the “each layer is an iteration” reading holds at initialization, and training replaces it with a measured spectral response. The rule is one line: learn the operator, compute the solver.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.