How Early Does Context Reduction Become Computationally Valuable? Capacity–Value Separation in Long-Context Prefill
Abstract
Long-context prefill is expensive because a large active context must be processed across many Transformer layers. Context reduction creates a fundamental tradeoff: earlier reduction lets more subsequent computation benefit from a shorter context, whereas deeper representations may support greater reliable reduction capacity. Yet early placement alone is insufficient: enough context must be removed reliably to yield a latency benefit. We ask whether the depth with the greatest reliable reduction capacity also minimizes measured prefill latency. Under a common reliability requirement, we determine at each depth the minimum token-retention ratio satisfying the same controlled evidence-retention threshold and measure prefill latency at that operating point. For Llama-3.1-8B and Ministral-8B, the capacity-optimal depths are L10 and L12, whereas the speedup-optimal depths occur earlier, at L2–L3 and L2. Crucially, these faster operating points retain 10–22 percentage points more context than their capacity-optimal counterparts yet achieve lower measured prefill latency. We call this mismatch capacity–value separation. Across the evaluated configurations, this separation persists when sufficient reliable reduction is available early and diminishes when reliable early reduction is limited. On LongBench, aggregate scores remain within 1.46 points of their downstream-only baselines across the evaluated early-layer grid, although category-level responses are more heterogeneous. These findings motivate treating the first context-reduction point as a conditional optimization decision, separate from later downstream selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.