ReFactor: Reflection-Augmented Agentic Video Generation
Abstract
Narrative video generation requires coordinating multiple tools and adapting subsequent operations to intermediate outputs. Existing approaches typically rely on fixed pipelines or one-shot plans, which often fail when generated outputs deviate from expectations and lack mechanisms for autonomous recovery. We propose ReFactor, a multi-turn agentic framework that formulates narrative video generation as a Markov decision process: At each step, the agent selects a tool call conditioned on the narrative prompt, execution history, and generated artifacts to maximize narrative completeness. Since video generation requires multiple sequential tool calls without informative feedback, the ability to reflection becomes crucial for agents to inspect generated artifacts, diagnose failures, and guide corrective actions. To enhance the agent's reflection ability, we construct 21K reflection-augmented trajectories grounded in actual tool executions and enriched with evaluator-accepted repair demonstrations for SFT. We further develop an agentic RL approach that introduces a reflection-anchored process reward to measure the relative quality improvements from intermediate repairs. This provides denser supervision for learning when and how to reflect across tasks of varying difficulties. On StoryEval, ReFactor achieves an 18.9% relative improvement in narrative completion over the strongest baseline. Beyond this gain, RL makes the policy learn selective reflection, cutting reflection frequency on easy prompts from 0.32 to 0.13 while raising it on hard prompts from 0.46 to 0.61 relative to the SFT policy, with scorer-defined repair success improving in both groups. Code is available at: https://anonymous.4open.science/r/Refactor-6BC2
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.