Long-Term Composed Video Retrieval: Benchmarking and Adaptive History Refocusing
Abstract
The rapid growth of video data has made Composed Video Retrieval (CVR) an emerging video retrieval paradigm. Existing CVR methods have achieved strong performance on short-term, i.e., single-turn, benchmarks, but they remain limited in real-world long-term retrieval. In practice, users often refine their requests iteratively based on intermediate retrieval results, rather than expressing the full intent in a single modification text. To systematically evaluate this challenge, we introduce LTCVR, a novel benchmark for Long-Term Composed Video Retrieval, featuring five representative history-dependency patterns: progressive refinement, rollback, cross-turn combination, local correction, and preference change. Our analysis on LTCVR shows that existing methods degrade markedly when the retrieval intent depends on earlier interaction context, and attention analysis further reveals that representative methods fail to focus on the key historical turns. To address this limitation, we propose InAHR, an Intent-Driven Adaptive History Refocusing framework. InAHR first infers the current retrieval intent and then selects a compact subset of historical visual evidence that is intent-relevant and representative. The selected evidence is re-injected into the generation process to produce retrieval-oriented target descriptions for video retrieval. Extensive experiments demonstrate that InAHR consistently outperforms strong baselines across diverse long-horizon CVR scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.