acceptodds
Under review as a conference paper at ICLR 2027

CORE-CVR: Benchmarking Fine-grained Reference Understanding in Composed Video Retrieval

Abstract

Composed Video Retrieval (CVR) retrieves a target video using a reference video and a modification text as a composed query. We reveal two issues with existing CVR benchmarks. First, they rarely require fine-grained understanding of the reference video. On three widely used benchmarks, 65.0–91.3% of the samples can still be retrieved when the composed query is replaced with a text of at most ten words. Second, they contain shortcuts that allow the target to be retrieved from a single query modality in 69.3–99.0% of the samples. Motivated by these findings, we introduce CORE-CVR, a new CVR benchmark in which both fine-grained reference information and the modification text are required to identify the target. CORE-CVR consists of Static-CVR, which preserves the fine-grained appearance of objects or backgrounds, and Temporal-CVR, which preserves the fine-grained execution of actions. We further introduce Reference Negatives and Modification Negatives, which can be rejected only by using the fine-grained reference information and the modification text, respectively. On CORE-CVR, only 22.1% and 16.6% of the samples are retrieved by short text and a single input, respectively. The best model achieves only 27.7% R@1, leaving substantial room for improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.