acceptodds
Under review as a conference paper at ICLR 2027

ProactGen: Learning Proactive Clarification for Video Generation

Abstract

Recent advances in video generation let users create high-quality videos from natural language descriptions, yet obtaining results that match user intent remains time-consuming: individual generations are computationally expensive, and underspecified prompts force repeated cycles of generation, inspection, and revision. Existing video prompt optimization methods improve generation through automatic rewriting, but cannot supply the ideas and preferences that were not in the initial description. We propose , a video generation framework that learns proactive clarification, eliciting user intent through questions grounded in generated video drafts. Specifically, we train a vision-language model as a questioning policy with reinforcement learning: observing the initial prompt and a video draft, it learns to ask the questions most valuable to downstream generation, rewarded by the alignment between the final generated video and a target reference. We further find that effective questioning does not require drafts of final-output quality, and exploit this property by generating drafts with fewer denoising steps, keeping the added cost of clarification low. Experiments show that improves target alignment over direct generation by at least the gain of the strongest baselines. Moreover, a single round of already exceeds the alignment that the strongest baseline reaches after four refinement rounds, in only 29% of its time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.