Benchmarking, Diagnosing, and Mitigating Temporal Positional Bias in Video Understanding
Abstract
Multimodal large language models (MLLMs) are sensitive to where relevant evidence appears in a video. Existing evaluations often manipulate evidence position by concatenating unrelated clips, which may introduce clip-matching shortcuts and disrupt the contextual relationships needed for reasoning. We investigate temporal positional bias through a unified study of benchmarking, diagnosis, and mitigation. We introduce TempPosBench, which shifts observation windows over continuous videos to vary evidence position while preserving temporal continuity. It comprises 1,322 question–answer instances across 13 tasks spanning perception, reasoning, and streaming awareness. Evaluating seven MLLMs reveals model- and task-specific positional preferences, with temporal reasoning exhibiting the strongest sensitivity. We further examine attention concentrated toward early or late video frames. Targeted prefix calibration and visual-key rotary position embedding (RoPE) compression provide evidence that attention sinks and positional-encoding-induced recency preferences can contribute to positional bias. For mitigation, we extend the benchmark construction to produce positionally balanced training candidates, then select positions that emphasize each model’s weaknesses. Fine-tuning on continuous video contexts improves both accuracy and positional robustness on spatial and temporal reasoning compared with concatenated clips. Under comparable training budgets, model-specific selection can also outperform training on all available positions. Fine-tuning additionally yields flatter frame-attention distributions, connecting the training-based mitigation to our attention-level diagnosis. Together, these findings highlight temporal continuity and model-specific positional supervision as complementary ingredients for mitigating temporal positional bias in video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.