Beyond Quadratic Attention: Separable Gaussian Operators for ODE Vision Transformers
Abstract
Self-attention carries a quadratic cost in token count, which limits how far Vision Transformers can scale with input resolution. We introduce SSoG, a separable sum-of-Gaussians spatial operator that replaces dense self-attention with a learned mixture of axis-factorized Gaussian kernels, reducing the per-layer cost from \(O(N^2)\) to subquadratic in \(N\). Unlike dot-product self-attention, whose Lipschitz constant is unbounded outside a compact input domain, SSoG with fixed kernels is globally Lipschitz continuous, giving the continuous-depth variant a well-posedness guarantee via the Picard–Lindel\"of theorem that dense-attention counterparts lack in general. These results establish SSoG as a practical, parameter-efficient, and theoretically grounded alternative to dense attention across discrete, weight-shared, and continuous-depth Transformer variants. On CIFAR-100 at \(224\times224\), in a matched discrete-Transformer setting and in the continuous setting, SSoG improves Top-1 accuracy by 2.22-4.12 points across three model widths while using 8.8-16.6% fewer parameters, with depth and training recipe held fixed. In a looped-Transformer setting, SSoG improves Top-1 accuracy by 5.4-8.14 points. We further probe inference-time scaling by evaluating on inputs up to \(1024\times1024\) at matched batch size: in the continuous-depth (Neural ODE) architecture, SSoG reaches up to 154% higher throughput and over 60% lower peak memory than dense attention, with the gap widening as resolution grows, consistent with the difference in asymptotic cost. In the discrete and looped settings, SSoG also achieves over 30% higher throughput than dense attention at this resolution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.