acceptodds
Under review as a conference paper at ICLR 2027

Triage Attention: Core-Protected Sparse Attention For Long-Video Vision-Language Models

Abstract

Attention accounts for a substantial share of prefill computation in long-video vision-language models. Sparse attention can reduce this cost, but inappropriate block selection may discard critical information. Existing selection strategies typically apply a uniform filtering rule without explicitly distinguishing the decision requirements of a high-confidence core from those of uncertain candidates. We introduce TRIAGE ATTENTION, a training-free sparse attention method that assigns distinct roles to two estimates of block importance. A head-specific score identifies a protected core and a larger candidate set, while a group-level score filters only candidates outside the core. This asymmetric rule preserves core support while restricting the auxiliary estimator’s authority to reject blocks, enabling a favorable quality cost trade-off. Across three backbones evaluated on the complete Video-MME and LongVideoBench benchmarks, Triage achieves accuracy close to dense attention. On Video-MME with Qwen3-VL-8B, it achieves 69.56% accuracy versus 69.44% for dense attention, while providing a 4.92×speedup in attention computation. A prespecified sweep of 15 configurations further demonstrates a competitive quality–cost trade-off against the evaluated ProxyAttn and XAttention configurations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.