acceptodds
Under review as a conference paper at ICLR 2027

TRAINING-FREE POSITIVE SPATIOTEMPORAL ENHANCEMENT FOR MITIGATING TEMPORAL HALLUCINATIONS IN VIDEO-LLMS

Abstract

Video Large Language Models (Video-LLMs) recognize objects and scenes reliably, yet they still misread event order, motion direction and object-state transitions, which produces *temporal hallucinations*. We refer to the change of the same object's state across frames as *object-centric motion*. Existing training-free approaches boost *per-frame* object-centric evidence to reduce hallucination. Yet, they do not explicitly boost *cross-frame* object-centric motions, the temporal relationships of the same object. Our empirical analysis suggests that this information is partially decodable from object-level visual representations yet, the final layer does not decode them in the final prediction. From this diagnosis, we propose Feature Enhance, a training-free scheme that strengthens object-centric visual evidence by emphasizing foreground regions and making inter-frame changes more explicit in the visual representations. It reinforces foreground regions that persist across frames and injects foreground-gated inter-frame feature residuals, leaving the model parameters untouched. Across 3 backbones, LLaVA-OneVision, Qwen2.5-VL and LLaVA-Video, Feature Enhance improves average accuracy over the corresponding unmodified model in all 6 backbone and benchmark combinations on VideoHallucer and EventHallusion, with gains of up to 4.67 and 4.00 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.