Same Features, Different Geometry: Localizing Muon's Advantage to the OV Circuit
Abstract
Does the optimizer change what a language model learns, or mainly how it stores and reaches it? We ask this at the level of top-k sparse-autoencoder (SAE) features: we compare Muon and AdamW on a 124M GPT-2-class model, matching SAE features across independently trained networks by firing-pattern correlation and calibrating against within-optimizer seed baselines, so every feature claim here concerns the firing-pattern content of one SAE family at one or two layers, not features in general. Under this protocol, cross-optimizer feature agreement (0.599) meets the AdamW seed ceiling (0.594), while Muon is more seed-reproducible (0.618 vs. 0.594) and changes weight geometry, a gap that is fully formed by one tenth of training, holds at three, six, and twelve heads, and narrows to the seed spread under a GeLU MLP; the loss advantage itself replicates from 30M to 1.3B parameters. Reassigning Muon between sub-circuits of a fixed architecture gives a double dissociation: Muon on every hidden matrix except the value circuit WV WO recovers none of the loss gain, Muon on the MLP blocks recovers none, and Muon on the query/key scoring circuit WQW ⊤K raises stable rank to near-full-Muon levels yet recovers neither loss nor reproducibility, whereas Muon on the value circuit alone recovers most of both, a sufficiency result that replicates a concurrent block- level ablation under a tuned baseline; zero-shot HellaSwag accuracy follows the same split. A rank-promoting AdamW control rules out endpoint stable rank as a sufficient knob, and varying only the fidelity of the orthogonalization step shows loss falling monotonically with fidelity while reproducibility saturates much earlier; a second matrix-aware optimizer, SOAP, reproduces the same signature. The opti- mizer invariance holds at three, six, and twelve heads, and matching features across layers shows that the metric moves when the compared features differ. In this bench, Muon’s useful interaction with attention runs through the value/output pathway and through orthogonalization fidelity, not through the scoring-rank fingerprint that observational analyses highlight.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.