Can Video Generation Agents Beat Simply Trying Again?
Abstract
A better video may come from a better prompt—or simply another try. Video generation agents revise prompts before generation or after inspecting results, yet they spend the same resource as best-of-N (BoN) sampling: generation attempts. Existing evaluations rarely compare agents against BoN with a simple prompt under the same budget, leaving unclear whether reported gains would survive such a comparison. We therefore ask: given the same number of attempts, can an agent beat simply trying again? We evaluate nine methods on a single reference-conditioned request across up to six video generators, giving each method 20 generations per generator and counting every video an agent produces, including those used only for feedback. Published agents are compared with Simple Prompt, which repeatedly samples a minimal prompt, and with a simple counterpart in their own family. The answer is largely no. No agent's best videos improve on Simple Prompt across all dimensions; gains in one dimension come with losses in another, especially in action fulfillment. More elaborate agents fare no better than Simple Prompt even when they improve on their simple counterparts, expert-written prompts perform on par, and the few gains rarely transfer across generators. Iterative agents also do not improve round by round: their final-round videos fall well below their best ones, and picking the best round turns them into BoN with a changing prompt. These results suggest that agents and simply trying again are far closer than published evaluations imply, and that agents should be evaluated against matched-budget BoN on the target generator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.