Post-Boundary Bridge: Must Local Attention Go Global Between Global Layers?
Abstract
Hybrid Transformers interleave local attention with periodic Full layers to cut long-context cost. Recent work on sparse and local attention reveals a gap between theoretical cross-layer reach and effective information propagation. This raises a key question: when periodic Full layers already provide global communication, how should local computation be allocated? We argue that local layers need not maximize cross-layer relay range; they can instead emphasize dense local modeling and timely nearby exchange. Post-Boundary Bridge (PBB) realizes this view by preserving dense causal attention within blocks while adding selective cross-block connections under a single softmax, with no extra parameters or routing. In controlled 200M experiments, PBB+Full matches SWA+Full in perplexity and improves the main source-retrieval probe; score-matched controls show that allocation, not only budget size, matters. In larger-scale transfer, PBB+Full stays close to Full attention in perplexity on 350M dense and 2B MoE models, and improves exact entity–attribute retrieval by 14.02 percentage points on the 2B MoE model. At 512K context, it achieves 3.32× prefill and 2.11× cached-decoding throughput over Full attention. These results support allocating local computation to targeted nearby exchange rather than maximum relay range.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.