Beyond the Final Reward: Depth-aware Post-Training for Instructional Video Editing
Abstract
Post-training with reinforcement learning is the standard route to aligning instruction-guided video editing models with human preferences. Current pipelines realize it by rolling out a full denoising trajectory and scoring only the final decoded video, typically with several specialized reward models whose scores are combined by hand. Such outcome-only supervision only observes the final output of each diffusion-transformer evaluation, treating the evaluation as a single atomic decision. In reality, nevertheless, the prediction is progressively assembled across dozens of instruction-conditioned blocks. Thus, this supervision cannot reveal at which depth editing quality is formed, nor which part of the instruction is responsible for it. To address this issue, we introduce a post-training method for **V**ideo editing through **I**ntermediate **S**elf-evaluation and **T**oken Attribution (**VISTA**). At selected transformer layers, VISTA first decodes groups of intermediate edit candidates, and evaluates them with a single reward model. A return-to-go objective then assigns each layer credit using its own reward and those observed at subsequent layers. The winning candidate is distilled into a token-importance memory that conditions later layers and re-weights a layer-wise GRPO update. By doing this, VISTA converts outcome-only GRPO into depth-aware self-evaluation without composing multiple specialized reward models. We further bound the bias caused by reusing the token-importance memory within a layer update, and give a sufficient condition under which this bias is dominated by genuine descent on the training objective. Experiments on 5B and 14B video-editing backbones demonstrate consistent improvements on OpenVE-Bench and FiVE-Bench, while ablations verify the complementary effects of intermediate self-evaluation and token attribution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.