acceptodds
Under review as a conference paper at ICLR 2027

SAW: Sink-Aware Weighting for Effective Video Reasoning

Abstract

Video reasoning requires linking evidence across time, yet providing frames spanning a video does not ensure that a multimodal large language model (MLLM) uses that evidence. We investigate frame-level attention sink, where response tokens concentrate attention on a small subset of frames. Our analysis shows that hallucinated tokens exhibit lower frame-attention entropy and higher peak attention than non-hallucinated tokens. We further derive an entropy-dependent bound on the single-frame approximation error of an attention head's visual contribution. Motivated by these findings, we propose Sink-Aware Weighting (SAW), which reduces the training influence of tokens with relatively concentrated frame attention. SAW converts frame-attention entropy into bounded token weights for supervised fine-tuning losses and reinforcement-learning advantages, while preserving the backbone architecture. With an 8B backbone, SAW achieves 64.99% average accuracy across five video question-answering benchmarks, 2.79 percentage points above the strongest evaluated open-weight baseline with complete results and comparable to GPT-4o (64.86%). On VideoHallucer, it improves pair-level accuracy by 7.60 percentage points over the strongest evaluated baseline. These results support attention-based reweighting for video reasoning and hallucination reduction. Data, code, and models will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.