acceptodds
Under review as a conference paper at ICLR 2027

COMPASS: Decentralized Language Model Post-training

Abstract

Pipeline-parallel post-training over Internet links must reduce communication while adapting open-weight models without architectural changes. The two-circuit design of anon2026twocircuit addresses this problem with a fast circuit that uses lossy compression in the forward and backward passes, and an anchor circuit that runs uncompressed passes on delayed weights. The anchor trades staleness for uncompressed computation and uses its gradients to guide the fast circuit. We propose COMPASS, which improves both pipeline compression and communication between the circuits. We propose to learn a basis from the anchor activations, reducing payloads at compressed boundaries by more than two orders of magnitude relative to dense communication in supervised fine-tuning (SFT). Across SFT and reinforcement learning with verifiable rewards (RLVR) experiments, we find that anchor-gradient signs are sufficient to provide useful guidance, reducing the cost of returning gradients. For RLVR, we propose a weight forecasting mechanism that accounts for anchor delay as the policy and its training data change. We also derive a convergence bound for the system and show that it achieves state-of-the-art results across various SFT and RLVR benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.