acceptodds
Under review as a conference paper at ICLR 2027

Granularity-Aligned Text-to-Video Retrieval with Semantic Conduits

Abstract

Text-to-video retrieval (T2VR) aims to align textual descriptions with video content, yet the core challenge lies in the inherent asymmetry of information density between the two modalities. Existing approaches often overlook this asymmetry in feature construction, either compressing videos into global representations or relying on coarse frame-level features, which conflate distinct visual concepts and yield a granularity mismatch with token-level textual features. To address this, we propose a granularity-aligned retrieval paradigm (GAR) that organizes video contents into semantic units better matched to fine-grained textual features. Drawing on the patch-level visual granularity that underpins fine-grained image-text alignment, we instantiate this paradigm with a hierarchical, joint spatial-temporal video representation built upon Semantic Conduits. Appearance-similar patches are first coalesced within each frame into semantic prototypes and then associated across frames. Soft patch-to-conduit assignments aggregate temporally coherent visual evidence into pooled conduit tokens, with a Transformer contextualizes by modeling relations among semantic components. MeanMaxSim aligns each text token with its most similar conduit and averages the token-wise scores, yielding fine-grained cross-modal correspondence while permitting video representations to be precomputed. Experiments on four benchmarks demonstrate strong retrieval performance. With a ViT-B/32 backbone, our method achieves R@1 scores of 48.4% on MSR-VTT, 48.1% on MSVD, 48.8% on DiDeMo, and 46.0% on ActivityNet. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.