acceptodds
Under review as a conference paper at ICLR 2027

Can we improve the performance of sparse attention in Grouped Query Experts without increasing experts?

Abstract

Sparse attention architectures, such as Grouped Query Experts (GQE), significantly reduce inference costs by routing input tokens to a small, statically fixed subset of query experts. However, rigidly discarding the unselected experts inevitably leads to information loss and degraded downstream performance. To address this, we introduce HyperGQE, a dynamic attention mechanism that recovers lost representational capacity by generating an auxiliary ”HyperExpert.” Rather than blindly adapting selected weights, HyperGQE computes a selection embedding based on the average representation of the unselected experts, which conditions a hypernetwork to generate dynamic bottleneck matrices (D and U). This HyperExpert processes the hidden states into a supplementary query projection (qhyper) that is integrated into a weighted attention slot alongside the standard routed experts. We evaluate HyperGQE on the SST-2 benchmark, observing that while standard GQE degrades accuracy to 75.92% compared to the dense GQA baseline (80.16%), HyperGQE not only fully recovers this gap but achieves 80.85% accuracy, exceeding the dense GQA baseline by 0.69 percentage points. Furthermore, hardware profiling demonstrates that HyperGQE achieves this superior accuracy using fewer active parameters (924.1M) than the dense baseline (946.0M), while maintaining competitive throughput and VRAM footprints across batch sizes. These results demonstrate that dynamically compensating for unselected experts allows sparse models to exceed the performance of dense counterparts without increasing the active parameter count beyond the dense baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.