Knowledge-Intensive Video Generation
Abstract
Text-to-video generation has advanced rapidly in visual quality, but existing task settings largely assume that the prompt already specifies what the video should depict. This assumption does not hold in information-seeking scenarios, where users typically specify what they want to learn rather than the detailed visual content of the answer. We introduce Knowledge-Intensive Video Generation (KIVI), where models must determine and generate appropriate visual content from concise information-seeking prompts that request explanations, procedures, or demonstrations. Many such requests naturally require long videos that organize multiple factual or procedural details over time. To study this capability, we construct KIVI-Bench, a benchmark of 1,080 prompts targeting long-form video generation, and propose automatic metrics for factuality and helpfulness. Human evaluation shows that our metrics achieve substantially stronger agreement with human judgments than existing visual-quality-oriented alternatives. Experiments on nine state-of-the-art video generation models reveal substantial variation in KIVI performance, while even the strongest models exhibit frequent failures in entity-specific visual properties, procedural operations, and clear information presentation. These results highlight the need to move beyond visually plausible rendering toward video generation systems that can reliably communicate correct and useful information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.