acceptodds
Under review as a conference paper at ICLR 2027

GOOSE: Which Video Does Not Belong? Multi-Video Reasoning for Out-of-Context Detection in Event Streams

Abstract

Recent multimodal large language models (MLLMs) show strong multi-video reasoning, yet social-media event streams pose a distinct challenge: authentic footage can still be *contextually false* when it originates from a different event. Existing benchmarks typically ask predefined questions about known targets, rather than requiring models to reconstruct an event and discover conflicting posts. To address this gap, we introduce **GOOSE** (**G**roup reasoning **O**ver **O**ut-of-context videos in **S**ocial-media **E**vents), a benchmark for discovering genuine footage falsely presented as evidence of another event. GOOSE contains 281 real-world cases, 351 OOC videos, and 4,085 videos across 6 domains, covering six evidence mechanisms involving scene, entity, temporal–causal, physical, multi-view, and provenance conflicts. It emphasizes **group-level evidence aggregation**, **event-level reasoning**, and **open-ended investigation**: models must discover unspecified OOC posts, while truthfully captioned off-event posts prevent relevance-based shortcuts. We evaluate two settings: Direct-IO jointly processes uniformly sampled frames, whereas a VideoAgent-style ReAct framework iteratively selects and inspects videos. Discovery is measured by case-macro OOC@K, and a rubric-guided native-video MLLM Judge evaluates grounded cross-video explanations; the Evidence-Grounded Score (EGS) jointly captures both aspects. Across 14 leading MLLMs, Gemini 3.8 Flash retrieves **86.73%** of gold OOCs but achieves only **58.56% EGS**, showing that correct discovery does not guarantee a valid explanation. VideoAgent also often trails Direct-IO because fragmented inspection weakens global event reconstruction and candidate–anchor comparison. These findings identify grounded cross-video verification, rather than visual plausibility or provenance recall alone, as the central bottleneck in real-world OOC detection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.