Mitigating Scene Transition Misunderstanding in Video LLMs via Semantic-Shift-Aware Contrastive Decoding
Abstract
Recognizing scene transitions is crucial for long-video understanding, yet Video Large Language Models (Video LLMs) often hallucinate nonexistent transitions or omit real ones. We analyze these failures through semantic-shift regions and find that over-attending to non-boundary shifts promotes hallucination, while under-attending to boundary-relevant shifts promotes omission. Controlled interventions provide evidence that such mis-calibrated attention causally contributes to scene-transition misunderstanding. Based on this finding, we propose SSACD (Semantic Shift-Aware Contrastive Decoding), a training-free method that preserves the original prediction in a clean branch while amplifying the identified attention tendency in an auxiliary branch and contrastively suppressing the resulting failure-prone predictions. SSACD realizes this principle through attention-level sharpening and input-level perturbation. Across multiple Video LLMs and benchmarks, SSACD outperforms existing training-free decoding methods on scene-transition understanding and improves multi-scene long-video performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.