TTPE: Test-Time Policy Evolution for Agentic Text-to-Video Refinement
Abstract
A single text-to-video (T2V) generation rarely satisfies every semantic requirement of a challenging prompt, so recent work spends additional video generations at test time on iterative generate–evaluate–refine loops. Existing methods refine each prompt in isolation under fixed strategies, repeatedly paying for video synthesis while discarding the experience accumulated from previous refinement episodes. We instead cast refinement as a budget-constrained sequential decision problem: each action shapes both the next video and the budget left for later correction. Conventional reinforcement learning is ill-suited here, as it requires costly video-generation rollouts and yields policies that must be retrained for each new generator. We introduce **TTPE (Test-Time Policy Evolution)**, a training-free framework that accumulates completed refinement trajectories in an episodic memory and reuses them to decide whether another generation is worth spending and which refinement action to take. From similar past trajectories, TTPE estimates two decoupled long-horizon advantages per action family: a quality advantage capturing expected future quality gain, and a cost advantage capturing expected future generation cost. Rather than collapsing them into a single objective, TTPE resolves their trade-off only at decision time, pruning Pareto-dominated actions in the quality–cost space and selecting among the survivors under the remaining budget. Across VBench-2.0 and StoryEval with three video generators, TTPE consistently improves the quality–cost frontier over prior iterative refinement, achieving the best aggregate quality in every evaluated generator–benchmark setting, with gains of up to 5.38 points and generation-cost reductions of up to 65.5% relative to fixed-round refinement. Code and implementation details are available at: https://anonymous.4open.science/r/ttpe-C245/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.