acceptodds
Under review as a conference paper at ICLR 2027

EditCast: One-Step Video Editing by Broadcasting an Edited Keyframe

Abstract

Image editors let users refine a frame’s appearance, but carrying edits through a video often requires another costly generation pass. We present EditCast, a one-step broadcaster that turns any image editor into a video editor. Given a source clip, an edited keyframe and its temporal slot, and an instruction, EditCast predicts the full edited clip in a single generator evaluation. We fully fine-tune VACE-1.3B using pixel and perceptual losses on sampled temporal windows. At equal training steps, a matched 30-step flow-matching editor fine-tuned from the same checkpoint performs worse on every keyframe-following and fidelity measure: by 1.7 dB at the keyframe slot, 2.0 dB at the last frame, and 2.2 dB against the reference clip. On KeyEdit-50, our curated 1280 × 720 benchmark with keyframes from four image editors and reference edits, EditCast is over 2 dB closer to the keyframe at the last frame than the strongest propagation baseline. It runs at 6.3 FPS on a single H200, or 16.8 FPS with TinyVAE, making it at least 33 times faster than all 50-step propagation baselines (87 times with TinyVAE). EditCast accepts one or more keyframes at any temporal slot, follows keyframes from different editors, and preserves 93% of the edit on clips 2.5 times the training length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.