Learning Iteration Schedules, Not Inverses: Zero-Shot Amortized Neural Preconditioning for Large-Scale Linear Systems
Abstract
Existing learned preconditioners are typically trained *per matrix*—with GNP requiring steps (s) per instance and GMP requiring epochs (s)—causing computational overhead to scale linearly with every solved system. Amortization fails because traditional approaches attempt to approximate the matrix inverse directly: a dense, domain-specific object of order . We shift the target of learning. Krylov solvers never consume the inverse; instead, they operate on a *configuration* comprising a bounded row scaler paired with step sizes for a degree- residual polynomial. This target is low-dimensional, scale-free, and universally defined across matrix families, enabling zero-shot prediction rather than per-instance optimization. The resulted proposed method, Mamba-PREC, distills a single network on a corpus disjoint from the benchmark. At inference, it sweeps matrix rows in Reverse Cuthill–McKee order via a bidirectional selective state-space model, predicts a configuration, validates it through a blind probe on a synthetic right-hand side, and freezes it for the solve. On the -matrix SuiteSparse benchmark, Mamba-PREC eliminates per-matrix training (s vs. s and s) with a median setup time of s, outperforms both baselines in time-to-tolerance and time-weighted AUC, and incurs zero non-finite breakdowns ( vs. and ). On the largest systems in the collection—every structurally nonsymmetric real SuiteSparse matrix of order above , the largest being —where GNP and GMP fail to construct on and instances respectively, Mamba-PREC dominates all four headline metrics at a median solve time of s. In pairwise comparisons on converged instances, Mamba-PREC reaches target tolerances faster in % of runs against GMP, % against GNP, % against classical AMG, % against Chebyshev semi-iteration, and % against thresholded GPU ILU, yielding median time-to-tolerance speedups of (GMP), (GNP), (AMG), (Chebyshev), and (ILU).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.