Count Is Not Enough: Inference-Time Mutation of Local–Global Attention
Abstract
Deployed hybrid attention models such as Gemma-3 interleave a small number of global-attention layers with sliding-window local layers in a fixed 5:1 ratio. We study how much of that trained pattern is actually load-bearing at inference time, by mutating layer types and window sizes of the frozen checkpoint — masking-only interventions that change no weights — and measuring exact negative log-likelihood (NLL), key–value cache memory, and decode throughput in one matched harness on A100-40GB GPUs. We re-implement NLL-guided fullattention layer selection, the strongest published training-free selection method, and compare it against uniform, positional, native-prefix, and periodic selection strategies at matched global-layer budgets. The result is a two-regime finding. At 1B parameters the trained pattern is near-optimal: NLL-guided selection improves on it by only −0.006 NLL (bootstrap 95% CI [−0.009, −0.004]), and cheap heuristics track the expensive baseline. At 4B parameters the trained pattern is beatable: NLL-guided selection finds a different set of five global layers that improves NLL by −0.085 (CI [−0.093, −0.077]) — the count of global layers is necessary but not sufficient, and which layers are global matters as models grow. A second finding is that the local window is a one-way knob: shrinking it from 512 to 128 tokens at 1B is free, while enlarging it past the trained value is catastrophic (NLL 3.66 →6.72 at 4k context). Converting global layers to local reduces KV-cache memory by up to 20× at 16k context with at most 2% decode-throughput cost at 1B, making the demoted configuration a practical long-context deployment on memory-bound cards.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.