PreCoG: Preview-Edit for Efficient Video Correction in Generation-Time
Abstract
Iterative text-to-video generation relies on feedback to correct compositional and temporal failures. However, existing methods typically require another full generation run for each correction, repeating computation and potentially disrupting already correct content. We introduce **PreCoG**, an iterative framework for correcting semantic failures during video generation. **PreQA** diagnoses failures from low-cost previews and constructs corrective conditions, while **PreEdit** edits cached generation trajectories under these conditions, attenuating orthogonal changes to limit disruption of already correct content.**PreCoG**completes generation only for the selected candidate, reusing its cached trajectory to reduce redundant computation and help preserve semantic consistency between the preview and the final video. Experiments show that **PreCoG** achieves the highest VBench score among the compared methods and approaches VSR's leading prompt alignment at its inference time. It also outperforms GENMAC on all three quality and alignment metrics at its inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.