STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models
Abstract
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STRAND (Spatio-Temporal Reasoning Audited by Necessary Decomposition), a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness. Crucially, we score models with Faithful Accuracy, an unconditional joint metric that credits a prediction only when the target answer and every prerequisite sub-question are correct, so that a model cannot inflate its score by being selectively consistent on the small subset of targets it happens to answer correctly. To address failure modes exposed by STRAND, we further propose STRAND-Track, an object-centric framework that explicitly constructs and reasons over structured object trajectories via chunk-wise state extraction and temporal aggregation. Extensive experiments, including comparisons against end-to-end MLLMs and modular video harnesses on a matched backbone, a configuration matched on frames, calls and tokens simultaneously, and a configuration matched on the answering prompt and its decoding settings, demonstrate that our object-centric framework significantly reduces hallucinated answers and improves spatio-temporal reasoning consistency over state-of-the-art MLLMs. Component ablations isolate cross-chunk identity linking as the step the gain rests on, while an oracle-trajectory reference shows that constructing the trajectories and reasoning over them leave comparable shares of the remaining error, so neither stage alone closes the gap.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.