acceptodds
Under review as a conference paper at ICLR 2027

Edited Endpoints as Partial Goals for Image-to-Video Generation

Abstract

Edited endpoint images make the target of an image-to-video transition explicit and can improve transition completion, but also preserve context in the source image's layout. Full-frame endpoint conditioning turns this preserved context into requirements on the video's final state, potentially inducing temporal artifacts. We measure them with Hard-Cut (HC) and Background Motion Reversal (BMR), for abrupt transitions and late background reversals, alongside Dynamic Degree (DD) for motion reduction, and observe model-specific artifacts across five video models spanning training-free and learned endpoint conditioning. We study Partial, a selective conditioning scheme that constrains only an edit-relevant region estimated from image-editor attention and leaves the remaining terminal content to the video model. Across three backbones on TC-Bench-I2V, Partial reduces temporal artifacts relative to full-frame conditioning, with no significant difference in Transition Completion Ratio (TCR); on Wan2.2-Turbo, uniformly weakening conditioning does not reproduce these reductions. A donor-context intervention, which replaces only the context outside the edit-relevant region with the final frame of a Partial video under the same full-frame conditioning, lowers HC and BMR on HunyuanVideo-1.5 and lowers BMR while restoring DD on Wan2.2-Turbo, with no significant difference in TCR, so constraint extent alone does not account for these artifacts. Together, these findings suggest that edited endpoints should serve as partial future goals, although edit relevance does not guarantee future validity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.