acceptodds
Under review as a conference paper at ICLR 2027

VideoRISE: Recursive Improvement through Self-Evaluation for Black-Box Video Inversion

Abstract

Recovering a prompt that reproduces a target video is difficult when a video generator is accessible only through queries: gradients and model internals are unavailable, each generation is stochastic, and semantic, appearance, motion, and temporal errors can improve or regress independently. We introduce VideoRISE (Recursive Improvement through Self-Evaluation), a budgeted control layer for black-box video prompt inversion. VideoRISE executes candidate prompts, evaluates target–output videos with a fixed multi-metric monitor, accepts a revision only after verification, and retains the best checkpoint together with failed-attempt evidence. Its reflective VLM proposer diagnoses differences in entities, action, style, camera, and lighting, then emits structured natural-language edits. On a 60-video, six-category protocol with Wan 2.2 Lightning T2V and I2V backends, the method improves the strongest one-shot VLM baseline from 0.607 to 0.683 in overall inversion score under 100 generations. We treat the score as a search and selection proxy rather than a perceptual ground truth, and report query, VLM-call, wall-clock, and estimated-cost accounting. Downstream prompt editing is shown qualitatively; broader multi-task use remains future work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.