acceptodds
Under review as a conference paper at ICLR 2027

TopoAtom: Query-Conditioned Semantic Atomization for Zero-Shot Video Temporal Grounding

Abstract

Zero-shot video temporal grounding (ZS-VTG) increasingly leverages frozen multimodal models to localize natural-language queries without task-specific training. Yet existing approaches largely build localization upon predefined temporal primitives, such as uniformly sampled frames or fixed-length clips, and subsequently measure their relevance to the query. Such fixed granularity overlooks a fundamental property of videos: semantics are structurally organized rather than uniformly distributed along the timeline. We introduce **TopoAtom**, a training-free framework that replaces predefined temporal primitives with adaptively discovered, query-conditioned semantic atoms. Based on pairwise clip distances, we build an -inspired connectivity filtration, whose component-merging process is compactly represented by a minimum spanning tree and its induced single-linkage hierarchy. Starting from the whole video, TopoAtom traverses this hierarchy top-down and recursively splits a component only when a finer descendant exhibits stronger query correspondence; otherwise, the component is retained as a semantic atom. These adaptive semantic atoms provide a structure-aware foundation for candidate generation. We further calibrate each candidate using a topology-derived persistence score, moving beyond sole reliance on query relevance by explicitly incorporating inherent connectivity and compactness. Experiments on Charades-STA, ActivityNet-Captions, and QVHighlights consistently demonstrate the effectiveness of TopoAtom.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.