acceptodds
Under review as a conference paper at ICLR 2027

RoboEdit: Reference-Guided Local Video Editing for Robotic Scenes

Abstract

Robotic manipulation videos are costly to collect and often tied to a specific embodiment. This embodiment-specific coupling limits reuse across robot platforms. We study reference-guided local video editing as a way to repurpose such demonstrations by replacing manipulators or objects while preserving the surrounding context. Since paired cross-embodiment videos are rarely available, we construct pseudo-paired supervision by extracting masks and reference images from the source video itself. However, this construction introduces a train-test mismatch: at training time, the reference is a crop of the source manipulator and is therefore tightly coupled to the foreground, whereas at inference, the reference is a novel image of a different manipulator. This mismatch can lead to reference leakage, boundary artifacts, and original-identity ghosting. We propose RoboEdit, a mask-aware consistency learning framework that localizes reference-induced changes to the editable region. RoboEdit compares reference-conditioned and null-reference forward passes and suppresses reference-induced changes in outside-mask token relations. We derive this objective from a flow-matching view of vector-field decoupling and implement it efficiently with token-level reservoir sampling. To evaluate this setting, we curate a robotic local editing benchmark and introduce Robot Manipulation Consistency (RMC), a task-specific metric for measuring local stability, trajectory preservation, and edited-region semantic consistency. Experiments show that RoboEdit improves reference fidelity, scene preservation, and manipulation consistency over existing video editing baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.