acceptodds
Under review as a conference paper at ICLR 2027

Dual Attention Residuals

Abstract

Recent work extends Transformer residual pathways along two complementary axes: historical retrieval provides direct access to earlier computation, whereas multi-stream methods maintain multiple evolving representations. These direc- tions have largely developed independently, leaving unexplored how depth-wise retrieval and multi-stream representation can complement each other within a uni- fied residual architecture. We propose Dual Attention Residuals (DAR), which combines these capabilities by allowing the two trajectories to participate recip- rocally in depth retrieval. For each target stream, DAR computes depth weights from normalized states in the opposite stream and applies them to values from the target stream’s own history. The retrieved states are combined as input to a Transformer branch. To control overhead, DAR retrieves over completed block histories while maintaining two temporary stream states within each block. Under matched training-token budgets, DAR achieves lower validation loss than the eval- uated residual baselines across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model. Parameter-matched ablations show that DAR benefits from both its dual-stream representation and reciprocal cross-stream retrieval. Rep- resentation and intervention analyses further show that reciprocal cross-stream selection preserves depth-wise diversity and avoids the redundancy or functional imbalance observed in alternative two-stream designs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.