Where You Put Attention Matters More Than How Much You Use: Attention Placement in Hybrid Sequence Models
Abstract
Hybrid sequence models use mostly recurrent blocks plus a few attention blocks, and most design studies ask how many attention blocks to use rather than where to put them. We compare six models, each with six blocks and the same number of parameters, so the only difference is where the attention blocks sit. We train each setup with several seeds on a single GPU and test on three tasks: parity state tracking, the A5 task (where order matters), and multi-query associative recall. On parity, putting one attention block at the front or in the middle beats both the recurrent baseline and back placement, while on recall the middle layout works well and the front layout almost fails. To find out why, we ran a paired test that removes attention after training: on parity, accuracy barely changes and the difference is not significant, but on recall it collapses and the difference is highly significant. This means the model needs attention at inference time for recall, while on parity attention mainly shapes what the recurrent blocks learn during training. Middle placement is the only layout that does well on both tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.