acceptodds
Under review as a conference paper at ICLR 2027

Structured Posterior Decoding over Contextualized Span Lattices for Video Moment Retrieval

Abstract

A central challenge in video moment retrieval is assigning reliable relevance scores among many plausible temporal spans. Recent work in audio moment retrieval showed that exact segment posteriors, even with simple boundary-based potentials, can provide an effective structured alternative to proposal confidence. We observe, however, that when this formulation is transferred directly to video, its gains diminish as more target moments must be retrieved. We attribute this to the different nature of temporal evidence in video: relevant visual semantics are often distributed throughout a moment, while boundary transitions tend to be less distinct. To address this, we contextualize the two-dimensional span lattice using query-conditioned local 2D convolutions, allowing each candidate to aggregate span-interior semantics and information from neighboring temporal hypotheses before structured inference. We then compute exact segment marginals and use them to form a posterior-weighted moment prototype, whose similarity to each candidate iteratively refines the span potentials. Under matched feature settings, our approach substantially improves multi-moment retrieval over prior methods on QV-M, and with InternVideo2-6B features it establishes a new state of the art on this benchmark. The gains carry over to QVHighlights and Charades-STA, while a cross-modal analysis on audio suggests that the benefit of span-lattice contextualization depends on modality-specific temporal evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.