acceptodds
Under review as a conference paper at ICLR 2027

ACT: Agentic Text–Scene Co-Design and Dual Evolution for Informative Video Generation

Abstract

Informative videos, such as tutorials, advertisements, and public-service videos, combine visual content and on-screen text to deliver specific messages. Generating such videos requires exact text to appear at the requested times and locations without obscuring the main subject. However, current methods often produce incorrect or poorly placed text and fail to coordinate it with the generated scene. To systematically study this problem, we introduce IVBench, a benchmark across five domains that jointly evaluates on-screen text, scene content and layout, and text-scene coordination. To improve text-scene coordination, we propose ACT, an agentic framework for text-scene co-design that plans text and scene together but renders them separately. ACT uses a shared layout contract to reserve space for text and guide subject placement before scene generation. It then adapts text placement to the generated scene, renders the required strings deterministically, and uses verification feedback for targeted repair. To reduce recurring errors and enable persistent system improvement, we further introduce verifier-guided dual evolution, which uses shared execution feedback to update both the harness and the video generator. Experiments on IVBench show that ACT consistently improves text control and text-scene coordination over strong commercial and open-weight baselines, with further gains from dual evolution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.