acceptodds
Under review as a conference paper at ICLR 2027

Breaking the Black Box: An Architecture-Specialized Preconditioning Optimizer for Attention Mechanism

Abstract

The increasing scale of deep learning models underscores the importance of efficient optimization. Existing matrix-based optimizers exploit parameter matrix structure through gradient statistics, activation statistics, or matrix transformations, but generally do not explicitly incorporate the analytical dependencies induced by specific network architectures. In this work, we propose STRATA, a structure-aware preconditioning optimizer for Transformer training. Moving from a black-box to a “gray-box” perspective, we explicitly exploit the coupling between the Query and Key projections in Attention to identify an architecture-specific structural curvature factor. We develop a two-stage approximation that transforms this activation-dependent quantity into a practical preconditioner constructed only from readily available model parameters and parameter gradients. We further characterize the approximation errors introduced by the two stages under a stylized model of the Attention distribution. Empirically, we evaluate STRATA on LLaMA-style pre-training from 130M to 1.3B parameters. Across the tested model scales, STRATA consistently improves optimization efficiency over AdamW and Muon and achieves end-to-end wall-clock speedups of 11%–31% relative to Muon, with negligible additional peak-memory overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.