acceptodds
Under review as a conference paper at ICLR 2027

Towards Recursive Self-Improvement: A Post-Training Paradigm for Video Generation

Abstract

Post-training has advanced large language models by converting feedback into effective supervision. Extending this process to video generation remains challenging: quality is multidimensional, automated evaluators imperfectly reflect human preferences, and aggregate scores provide limited guidance on what to improve next. We introduce a closed-loop, backbone-agnostic post-training paradigm that connects capability-level evaluation to targeted interventions in data collection and training. The paradigm couples data and training loops through human-calibrated evaluation, combining task-specific metrics, order-controlled pairwise comparisons, and statistically grounded checkpoint promotion criteria. Experimental outcomes guide revisions to training recipes, data mixtures, and collection priorities, while validated improvements are carried forward into subsequent iterations. This feedback-to-intervention mechanism enables adaptation of both the generator and the process used to produce its successors, offering a practical path toward system-level recursive self-improvement. We instantiate the paradigm on several open-weight and licensed video foundation models, including LTX-2.3/2.5, Wan-2.2, and MiniMax-H3. Post-trained LTX checkpoints win 88.2% and 80.6% of decisive human comparisons; Wan-2.2 reduces audio–visual offset from 6.25 to 1.69 frames. H3 achieves an 85.3% advertising reference-fidelity pass rate, the highest among four models and six frontier video-generation APIs, while remaining competitive on general video quality. These results demonstrate the applicability of the paradigm across model families and provide a foundation for studying recursive self-improvement in video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.