Inducing Hallucinations in Video Large Language Models by Exploiting Spatiotemporal Corruption Sensitivity
Abstract
Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in open-ended video understanding, yet remain susceptible to video manipulations. Existing video manipulations mainly aim to mislead Video-LLMs into producing wrong labels or predefined responses, leaving hallucinations, i.e., plausible responses containing multiple visually unsupported errors, unexplored. In this paper, we address this research gap by proposing VidHaze, the first approach to inducing hallucinations in Video-LLMs via video perturbations. Our key observation is a clear sensitivity gap between correct and hallucinated responses to diverse spatiotemporal video corruptions, where correct responses exhibit larger clean-to-corrupted distribution shifts. Building on this observation, VidHaze selects spatiotemporally corrupted videos as hard negatives and contrasts their next-token distributions with those of clean videos under a shared prefix, converting the resulting directional shifts into hallucination-oriented soft targets. Experiments on VATEX and MSRVTT datasets across four advanced Video-LLMs show that VidHaze achieves average sentence-level hallucination rates of 60.67%/34.27% in white-/black-box settings, while maintaining clean-response linguistic quality. Further analyses on when and why VidHaze is effective show its dependence on model family and dataset characteristics, and that its gains stem from increased hallucination density rather than response length.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.