RadFree: General Depth-Aware Orthogonal Residual Writes for LLMs
Abstract
Standard residual connections in LLMs add entire sublayer outputs to the residual stream. We show that the radial components of these output parallel to the residual state contribute little to language modeling in deep layers, yet drive substantial residual growth. We introduce Radial-Free Residual Write (RadFree), a simple update that selectively removes radial components in late layers using a detached projection reference while preserving perpendicular updates. We evaluate RadFree at frontier scale through modern MoE pre-training up to 125B total parameters. RadFree consistently lowers validation loss and improves performance on nearly all downstream benchmarks, raising average accuracy by approximately 1.5 percentage points across scales. Combined with manifold-constrained hyper-connections (mHC), RadFree improves average accuracy by 2.93 percentage points over the standard residual baseline, more than doubling the 1.29-point gain of mHC alone. RadFree also makes residual growth in deep layers more gradual, reducing final residual RMS by 36–37% across scales. RadFree adds no parameters and achieves zero training and inference overhead through simple kernel fusion, offering a simple and practical way to improve LLM performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.