acceptodds
Under review as a conference paper at ICLR 2027

Modeling Retrieval Intent from Video-Text Context for Composed Video Retrieval

Abstract

Composed Video Retrieval (CVR) retrieves a target video by composing a reference video with a text modification describing desired changes. Effective CVR requires resolving the user's retrieval intent from joint video-text context: what to modify and what to preserve. However, existing methods merely perform task-specific refinement within each modality, underexploring how video-text interaction determines user's retrieval intent. This leads to two critical weaknesses: 1) imprecise modification, where changes are not confined to relevant visual regions, and 2) indiscriminate preservation, where insignificant video content is fully retained. To address this, we propose IC-CoVR, an intent-centric framework that resolves the retrieval intent within the joint video-text context. Specifically, we characterize retrieval intent along two complementary dimensions: explicit intent for targeted modification and implicit intent for selective preservation. Based on this formulation, IC-CoVR first constructs structured video representations to support video-text interaction and intent reasoning, then uses explicit intent to locate and guide targeted modification, and finally leverages implicit intent to selectively preserve unmodified content by suppressing irrelevant details. Extensive experiments on FineCVR-1M and EgoCVR show that IC-CoVR achieves state-of-the-art performance and strong generalization, with a substantial gain of +9.22% R@1 on FineCVR-1M. Comprehensive ablation studies further validate the effectiveness of our intent-centric formulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.