TopoAtom: Query-Conditioned Semantic Atomization for Zero-Shot Video Temporal Grounding
Abstract
Zero-shot video temporal grounding (ZS-VTG) increasingly leverages frozen multimodal models to localize natural-language queries without task-specific training. Yet existing approaches largely build localization upon predefined temporal primitives, such as uniformly sampled frames or fixed-length clips, and subsequently measure their relevance to the query. Such fixed granularity overlooks a fundamental property of videos: semantics are structurally organized rather than uniformly distributed along the timeline. We introduce **TopoAtom**, a training-free framework that replaces predefined temporal primitives with adaptively discovered, query-conditioned semantic atoms. Based on pairwise clip distances, we build an -inspired connectivity filtration, whose component-merging process is compactly represented by a minimum spanning tree and its induced single-linkage hierarchy. Starting from the whole video, TopoAtom traverses this hierarchy top-down and recursively splits a component only when a finer descendant exhibits stronger query correspondence; otherwise, the component is retained as a semantic atom. These adaptive semantic atoms provide a structure-aware foundation for candidate generation. We further calibrate each candidate using a topology-derived persistence score, moving beyond sole reliance on query relevance by explicitly incorporating inherent connectivity and compactness. Experiments on Charades-STA, ActivityNet-Captions, and QVHighlights consistently demonstrate the effectiveness of TopoAtom.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.