ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Abstract
Multi-reference video editing integrates content from multiple reference images while preserving relevant content from a source video. However, operation-centric instructions often leave preservation requirements and content-source relationships underspecified. Even detailed instructions may associate entities or attributes with the wrong visual sources, introducing inconsistent training supervision. To address these limitations, we introduce ReBind, a supervision framework that combines comprehensive editing and preservation semantics with accurate source attribution. As it’s core, ReBind-Instruct generates comprehensive, reference-consistent instructions through a two-stage training process. To learn these capabilities, Structured Instruction Learning establishes instruction generation through supervised fine-tuning, while Reference Attribution Optimization applies GRPO with content-level and relation-level rewards to improve semantic completeness and correctness and penalize source assignment errors. The resulting annotations provide scalable supervision for ReBind-Edit, which extends a pretrained text-to-video model to multi-reference video editing through joint visual conditioning on the source video and reference images. Controlled ablations show that corrupting source assignments degrades editing performance despite largely retained instruction quality, demonstrating the value of accurate attribution. ReBind-Instruct achieves the highest combined Instruction Quality and Reference Accuracy score among evaluated MLLMs, while ReBind-Edit achieves state-of-the-art open-source performance on IntelligentVBench and UniVBench across single- and multi-reference editing. We will publicly release both models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.