acceptodds
Under review as a conference paper at ICLR 2027

MORSE: Learning Component-Wise Residual Transformations for Composed Video Retrieval

Abstract

Composed Video Retrieval (CoVR) retrieves a target video given a reference video and a text instruction that describes the change. Existing methods learn composed representations through target alignment, but global target alignment alone is insufficient to distinguish the intended target from visually similar candidates that fail to satisfy the requested modification. How to coordinate changes across the components of the reference representation to capture the instruction remains underexplored. We propose MORSE, which formulates target-feature prediction as context-conditioned corrections over an ordered residual decomposition. This decomposition organizes a reference feature into additive components with cumulative dependencies. Residual Transformation predicts differentiated corrections conditioned on the instruction and each component’s residual context, then aggregates them with the reference feature to form a target-feature prediction. Target-supervised Refinement constrains the cumulative composition of the transformed components using the target decomposition under shared codebooks, connecting overall target discrimination with structural supervision without intermediate change annotations. Compared with recent methods on each benchmark, MORSE achieves relative improvements of 17.81% in R@1 on FineCVR-1M and 10.25% in mAP@50 on TF-CoVR. Further analyses support the effectiveness of context-conditioned corrections and target-side structural supervision in distinguishing instruction-specific changes. Code will be available upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.