Read Depth and Broadcast Depth: What Sharing Keys and Values Across Layers Costs a Transformer
Abstract
Language models increasingly share keys and values across layers, so many layers query caches that few layers compute. We separate the two roles of attention: the read depth ρ counts layers that query, or select from, the whole context afresh, and the broadcast depth β counts the key–value generations that global attention reads, each computed below its readers. Sharing preserves the first and not the second. The composition construction of Chen, Peng and Wu (FOCS 2025) works with one generation, and reusing or re-ranking selections within a generation adds no read depth: for a fixed architecture at logarithmic precision, read depth ρ cannot compose ρ functions on long enough prompts whose blocks lie beyond the sliding windows' reach. With one generation of per-token keys and selections and the windows below it, the hierarchy holds at linear length: following ρ+1 pointers between two m-entry tables farther apart than the windows' reach needs read bandwidth (the bits one side of the prompt sends over all reads) Ω(m) − O(ρ log m), less one hidden state, at read depth ρ, but O(ρ log n) at read depth ρ+1. Without sliding windows, layers reading one such generation need read bandwidth N/2 − log₂(1/(2ε)) to compute the inner product of two N-bit blocks with advantage ε over uniform inputs, whereas two generations compute it with two layers of logarithmic size. In these terms YOCO with a sliding-window self-decoder has broadcast depth 1 and read depth L/2, its sparse variant YOIO has read depth 1, and DeepSeek-V4.1-Flash, as specified, has broadcast and read depth 4 among 38 global layers; at published widths and p = 16 the bounds for YOCO's joins and YOIO's two-pointer chasing apply from prompts of order 10^7 tokens, but DeepSeek-V4.1-Flash's composition bound needs over 10^100000 tokens. Bounds concern one forward pass in p-bit arithmetic with exact attention partial sums, p possibly growing with the prompt length n.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.