Rethinking Hallucination in Long-Video Understanding: A Unified Binding Perspective and a Benchmark
Abstract
Hallucination is a major bottleneck in reliable long-video understanding. Unlike short clips, long videos involve expanded context windows, dense temporal dynamics, and complex multi-entity narratives. These factors make long-video hallucination far more complex and diverse. Existing methods tackle this challenge in a piecemeal fashion, relying on task-specific heuristics, such as frame selection or evidence search, to mitigate individual hallucination types. We rethink long video hallucination, arguing that these diverse hallucination stem from a single root cause: incorrect multimodal binding. We establish a unified perspective and a comprehensive taxonomy dividing atomic errors into two families: out-video hallucination (binding to non-existent entities) and in-video hallucination (misbinding valid elements to wrong entities, roles, states, or timing). To address this, we introduce a temporal hypergraph framework for correct-by-construction understanding without task-level training, using n-ary relations and temporal validity to eliminate fabrication and misattribution. We further build a taxonomy-labeled benchmark for long video hallucination that not only measures the overall hallucination level of existing models and frameworks, but also diagnoses the source and the bias of hallucination, pointing out directions for improvement. Across many benchmarks, our approach yields consistent gains over strong baselines. This work offers a unified paradigm to tackle long video hallucination, moving beyond fragmented, task-specific heuristics. Code and data are available at https://anonymous.4open.science/r/rrt_echo-874C/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.