acceptodds
Under review as a conference paper at ICLR 2027

GEEZ: Grounded, Efficient, and Effective Video Understanding

Abstract

Existing video benchmarks can contain unanswerable or subjective questions, or questions solvable without the intended audio-visual evidence. We present GEEK, a benchmark of 1000 multiple-choice questions on web videos lasting 15 s. to 5 min.: 800 answerable from the video alone and 200 linking video evidence to verified external knowledge. We use a multimodal large language model to build EntityScript, a timestamped record of entities, events, and audio cues. An optimizer model iteratively refines prompts for generating questions from EntityScript, while web-search agents construct questions involving external knowledge. Candidate screening combines adversarial filtering with manual review. We also present GEEZ, a video question-answering framework that augments EntityScript with speech transcripts and on-screen text extracted offline. A planner prioritizes the transcripts for spoken wording and the extracted text for on-screen content, while querying an audio-visual recognizer for visual details. With Qwen3-Omni-30B as the recognizer, GEEX achieves 67.4% accuracy on GEEK. Compared with the EntityScript-based agent without these additional text sources (ES-AGENT), it improves accuracy by 8.5 percentage points while using 24% fewer query-time recognizer calls. It also outperforms a planner using only EntityScript, speech transcripts, and on-screen text by 17.9 points. Across recognizers, gains over ES-AGENT range from 4.1 to 5.1 percentage points on OmniVideoBench and Video-MME-v2. We will release the benchmark, construction records, video annotations, and framework code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.