acceptodds
Under review as a conference paper at ICLR 2027

ROLE-DECOUPLED ATTENTION RESIDUALS: SEPARATE QK AND V READS IN BLOCK ATTNRES

Abstract

Block Attention Residuals use one content-dependent depth mixture to construct queries, keys, and values. This couples the representations that determine attention scores to the value content aggregated under those scores. Role-Decoupled Attention Residuals (RD-AttnRes) retain the parent route and learn a second route for over the same Block sources. Tying the route queries recovers Block AttnRes; untying them adds one model-width vector per attention layer and no second token-to-token attention operation. After 2B FineWeb-Edu tokens, RD-AttnRes lowers fixed-prefix validation NLL in all five paired seeds at both 120M and 343M parameters, with mean within-seed perplexity reductions of 2.97% and 2.43%. In a single-seed 1.5B pair trained for 10B tokens, RD-AttnRes lowers endpoint perplexity by 2.22%, leads at all 20 shared quick-validation checkpoints, and reaches the same validation NLL with 8.89% fewer training tokens on average across 17 interpolation-supported targets. The four primary zero-shot scores improve by 1.77 percentage points on average. Tested controls and fixed-checkpoint interventions show that trained models learn and use the separation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.