Dobb: Shard-Local Preconditioning in Block-Diagonal Eigenbases
Abstract
Training language models whose parameters and optimizer state exceed a single node's memory requires sharding these tensors across devices. Under tensor parallelism (TP), matrix-aware optimizers such as Muon introduce communication to perform global matrix operations across shards. We propose Dobb (Distributed Optimizer with Block-diagonal Bases), a matrix-aware optimizer designed to compute updates locally within parameter shards. Dobb approximates the input-activation second moment with diagonal blocks and tracks a complete basis within each block. It projects momentum into these bases, applies sign normalization to the coefficients, and reconstructs the update, retaining every direction within each block. When each block lies within one shard under TP or fully sharded data parallelism (FSDP), Dobb applies its matrix update locally, requiring no communication beyond data-parallel synchronization of the activation second moments. We benchmark Dobb against Dion, a distributed matrix-aware optimizer designed to retain the benefits of orthonormalized updates while reducing communication. On GPT-2 models with 124M, 354M, and 772M parameters, Dobb with 64 blocks achieves validation losses of 3.037, 2.742, and 2.600, respectively, compared with 3.265, 2.879, and 2.696 for Dion at rank fraction 1/64. In a Qwen3-32B timing benchmark with FSDP shard degree 32 and TP degree 2, Dobb reduces the amortized optimizer-call time from Dion's 1307.7 ms to 168.0 ms, a 7.8 speedup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.