Dual Attention Residuals
Abstract
Recent work extends Transformer residual pathways along two complementary axes: historical retrieval provides direct access to earlier computation, whereas multi-stream methods maintain multiple evolving representations. These direc- tions have largely developed independently, leaving unexplored how depth-wise retrieval and multi-stream representation can complement each other within a uni- fied residual architecture. We propose Dual Attention Residuals (DAR), which combines these capabilities by allowing the two trajectories to participate recip- rocally in depth retrieval. For each target stream, DAR computes depth weights from normalized states in the opposite stream and applies them to values from the target stream’s own history. The retrieved states are combined as input to a Transformer branch. To control overhead, DAR retrieves over completed block histories while maintaining two temporary stream states within each block. Under matched training-token budgets, DAR achieves lower validation loss than the eval- uated residual baselines across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model. Parameter-matched ablations show that DAR benefits from both its dual-stream representation and reciprocal cross-stream retrieval. Rep- resentation and intervention analyses further show that reciprocal cross-stream selection preserves depth-wise diversity and avoids the redundancy or functional imbalance observed in alternative two-stream designs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.