SITE: Surfacing Inter-frame Temporal Evidence to Recover VLMs' Dormant Temporal Reasoning
Abstract
Vision-Language Models (VLMs) typically answer temporally discriminative video questions using uniformly sampled frames. This interface is efficient, but presents the same frame-wise input regardless of the temporal relation required by the question. Even when the frames contain the relevant visual observations, these relations remain implicit. Answering correctly therefore requires the VLM to both organize these observations into usable temporal evidence and reason over that evidence. These two operations are coupled within a single prediction, so final-answer accuracy alone cannot distinguish failures in evidence formation from failures in answer reasoning. We investigate whether explicitly presenting question-relevant temporal evidence can correct errors in frozen-VLM under this interface. We introduce SITE, a training-free Route–Surface–Corroborate framework. SITE routes each question to the required type of inter-frame evidence, constructs corresponding evidence views through visual organization and task-specific instructions, and combines their predictions through cross-view voting. Across three benchmarks and three frozen backbones, SITE improves accuracy by up to 23.0 percentage points over the sampled-frames baseline and outperforms matched-backbone training-free baselines. These results show that explicitly organizing relevant temporal evidence and corroborating predictions across views can correct temporal QA errors without updating model parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.