acceptodds
Under review as a conference paper at ICLR 2027

GroundedVDR: Grounding Video-to-Web Evidence Dependencies in Video Deep Research

Abstract

Video large multimodal models have improved at identifying cross-frame visual evidence and conducting open-web retrieval, as reflected in recent video deep research (VDR) benchmarks. However, existing VDR benchmarks primarily evaluate tool-use trajectories and final responses without verifying whether models use temporal video information or retrieved web evidence. We introduce GroundedVDR, an evidence-grounded benchmark with three features. First, it combines human review and model testing to exclude questions that can be answered using parametric knowledge alone. Second, it provides structured video-to-web evidence annotation that links related video frames with their corresponding web evidence, while recording candidate-linked hard negatives for diagnostic analysis. Third, it measures video-state localization, evidence retrieval, video-to-web evidence alignment, and final answer synthesis separately. Together, these features allow GroundedVDR to evaluate both answer correctness and support from the required video and web evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.