acceptodds
Under review as a conference paper at ICLR 2027

Exact Where It Matters: Downstream-Guided Attention Refinement for Long-Video Understanding

Abstract

Approximating attention accelerates long-video understanding, but local error alone does not determine whether a correction improves the answer. We introduce EWiM (Exact Where it Matters), a training-free method for downstream-guided attention refinement during offline video–question prefilling. At selected layers, EWiM probes a bounded candidate pool through isolated corrections on sampled query rows and measures their effects on a question-conditioned next-layer readout at the terminal prompt position. Incremental updates reuse the common coarse state and refresh only the changed contributions to the next-layer readout. An execution-aware allocator uses these one-step sensitivity scores and hardware-profiled marginal costs to select full-block refinements, updating costs to account for shared row computation within a budget that includes probing overhead. Input tokens, model parameters, and the decoding procedure remain unchanged. Across eight main model–benchmark pairs, EWiM retains 99.2% of dense accuracy on average at a 1.46× geometric-mean speedup in time to first token, including visual encoding and all acceleration overhead. Controlled interventions show that the selected block sets yield higher signed answer-margin utility than those chosen by local-error controls. Further experiments demonstrate transfer to additional backbones and composition with token compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.