Accelerating Mixture of Block Attention via Dynamic Routing, KV-Centric Execution and Runtime Pipelining
Abstract
Mixture of Block Attention (MoBA) reduces long-context computation by selecting key-value (KV) blocks for each query. FlashMoBA implements MoBA efficiently but retains a fixed Top- budget and query-centric schedule. Fixed Top- applies one budget across different score distributions and couples routing to task granularity; query-range ownership fragments K/V reuse when several ranges select the same column. We present LightningMoBA, which co-designs selection and execution through dynamic routing, KV-centric ownership, and runtime pipelining. Dynamic routing thresholds MoBA's per-query block scores and directly emits variable-cardinality sparse metadata. KV-centric ownership groups selected rows by (query head, KV block) column, splits high-fanout columns into bounded tasks, and reuses each task's loaded KV tile. On Hopper, pipelining overlaps regular K/V transfers and irregular query gathers with Tensor Core computation. Together these mechanisms decouple retention from GPU scheduling granularity. At matched Top- retention across five checkpoints, LightningMoBA accelerates end-to-end generation by to while raising mean RULER score from to . In controlled two-model replays, complete device-path speedup grows from to at Top- to to at Top-, while routing and CSC construction reach to at Top-. The code of LightningMoBA is available at https://anonymous.4open.science/r/LightningMoba-A090
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.