acceptodds
Under review as a conference paper at ICLR 2027

Grounding Confusable Movie Moments

Abstract

Characters, locations, and actions recur throughout a movie, so many of its moments differ only in fine-grained details. Grounding such a moment in a full-length movie requires resolving that detail rather than matching the attributes it shares with other moments, yet existing temporal movie grounding benchmarks rarely test this. We introduce CoM-Bench, a benchmark of confusable query pairs, where the two queries in each pair share most attributes but differ in a single distinguishing detail within a full movie, posing a challenge for recent temporal grounding approaches. To locate where grounding breaks down, we analyze a retrieve-then-localize pipeline, in which a retriever first selects a short candidate chunk and a localizer then performs temporal grounding within it. Separating the two stages lets us attribute each failure to either retrieval or localization. We find that retrieval is the major bottleneck: state-of-the-art retrievers often select a moment's confusable counterpart, because their choices follow the content the two moments share rather than the detail that tells them apart. To address these failures, we propose a simple, training-free remedy, Agentic Window Search (AWS): an MLLM agent the query into attribute-level sub-queries, then each candidate window against them and the window when evidence is cut off at its edge. AWS works with existing retrievers and localizers without additional training, raising mIoU from 18.5 to 26.8.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.