acceptodds
Under review as a conference paper at ICLR 2027

Inducing Hallucinations in Video Large Language Models by Exploiting Spatiotemporal Corruption Sensitivity

Abstract

Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in open-ended video understanding, yet remain susceptible to video manipulations. Existing video manipulations mainly aim to mislead Video-LLMs into producing wrong labels or predefined responses, leaving hallucinations, i.e., plausible responses containing multiple visually unsupported errors, unexplored. In this paper, we address this research gap by proposing VidHaze, the first approach to inducing hallucinations in Video-LLMs via video perturbations. Our key observation is a clear sensitivity gap between correct and hallucinated responses to diverse spatiotemporal video corruptions, where correct responses exhibit larger clean-to-corrupted distribution shifts. Building on this observation, VidHaze selects spatiotemporally corrupted videos as hard negatives and contrasts their next-token distributions with those of clean videos under a shared prefix, converting the resulting directional shifts into hallucination-oriented soft targets. Experiments on VATEX and MSRVTT datasets across four advanced Video-LLMs show that VidHaze achieves average sentence-level hallucination rates of 60.67%/34.27% in white-/black-box settings, while maintaining clean-response linguistic quality. Further analyses on when and why VidHaze is effective show its dependence on model family and dataset characteristics, and that its gains stem from increased hallucination density rather than response length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.