ACT: Agentic Text–Scene Co-Design and Dual Evolution for Informative Video Generation
Abstract
Informative videos, such as tutorials, advertisements, and public-service videos, combine visual content and on-screen text to deliver specific messages. Generating such videos requires exact text to appear at the requested times and locations without obscuring the main subject. However, current methods often produce incorrect or poorly placed text and fail to coordinate it with the generated scene. To systematically study this problem, we introduce IVBench, a benchmark across five domains that jointly evaluates on-screen text, scene content and layout, and text-scene coordination. To improve text-scene coordination, we propose ACT, an agentic framework for text-scene co-design that plans text and scene together but renders them separately. ACT uses a shared layout contract to reserve space for text and guide subject placement before scene generation. It then adapts text placement to the generated scene, renders the required strings deterministically, and uses verification feedback for targeted repair. To reduce recurring errors and enable persistent system improvement, we further introduce verifier-guided dual evolution, which uses shared execution feedback to update both the harness and the video generator. Experiments on IVBench show that ACT consistently improves text control and text-scene coordination over strong commercial and open-weight baselines, with further gains from dual evolution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.