FlowSMO: Scalable End-to-End Sparse Matrix Operators via Latent Flow Matching
Abstract
Large-scale scientific computing often requires repeatedly applying matrix operators associated with sparse matrices , including inverse and inverse-square-root operators. Classical algebraic methods, including incomplete factorizations, Krylov subspace methods, and polynomial or rational approximations, remain effective but are difficult to map efficiently to modern GPUs due to problem-specific iterations, triangular solves, and global synchronization. Learning-based methods offer a promising alternative, yet existing approaches are largely confined to inverse operators for linear system solving and often require CPU-side refinement or non-GPU-resident post-processing, breaking end-to-end GPU execution. We propose FlowSMO (Flow-based Sparse Matrix Operator), a size-generalizable GPU-resident generative framework for approximating sparse matrix operators across varying matrix dimensions. For a given operator class, FlowSMO first encodes the input sparse matrix into a dimension-agnostic conditioning representation. Conditioned on this representation, a flow-matching model samples latent operator variables with a cost independent of the matrix size. To recover an operator at the original scale, a sparse decoder predicts a local sparse component over candidate edges induced by the input sparsity pattern, while a low-rank component captures global correlations and fill-in effects. The resulting operator is represented implicitly as , enabling matrix-free application without explicitly assembling dense matrices. All stages, including generation and refinement, are executed on GPU, with computational cost scaling as for rank . We validate FlowSMO across three complementary operator settings: Helmholtz inverse operators for preconditioner generation under moderate spectral difficulty, inverse-square-root operators on sparse SPD matrices, and Schur complements as block-structured operators beyond standard spectral matrix functions. Experiments show that FlowSMO improves approximation quality while enabling GPU-resident near-linear inference and cross-scale generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.