ACTIV: Reducing Spatiotemporal Hallucinations in Video Language Models through Token Interventions
Abstract
Video language models have gained adoption in applications such as robotics and autonomous driving, where reliability and low inference latency are essential. Yet, these models still misreport where objects are and the order in which events occur. Contrastive decoding can reduce these errors without retraining by comparing predictions from original and altered visual inputs. However, constructing negatives by re-encoding altered videos increases inference latency. We target a useful accuracy–latency tradeoff by constructing spatial and temporal negatives directly from already-computed visual tokens. We introduce Adaptive Contrast with Token Interventions for Video (ACTIV), a training-free framework that constructs spatial and temporal views with only a single video encoding pass. ACTIV averages visual features across locations or time to bottleneck the corresponding information. During answer generation, ACTIV contrasts predictions from these alternatives with those from the original features, weighting each comparison by its effect on the prediction to favor words supported more strongly by the original video. ACTIV requires a single visual encoding and three decoder branches, achieving competitive hallucination mitigation with substantially lower latency than re-encoding-based contrastive decoding. Across three models and three hallucination benchmarks, ACTIV outperforms the evaluated training-free baselines in average score for each model; on LLaVA-Video-7B, it improves VidHalluc and EventHallusion scores by 3.87 and 5.38 points over the unmodified model, respectively, while reducing estimated model-inference latency by approximately **22–26**% relative to the state-of-the-art.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.