acceptodds
Under review as a conference paper at ICLR 2027

DecoupleGen: Intent-Aware CoT Layout Planning and Controllable Text Rendering for Video Covers

Abstract

Video covers are often a viewer's first encounter with a video, yet their typography must be composed over a content-bearing frame whose people, objects, actions, and native screen content remain recognizable. We study this progressive design problem with DecoupleGen, which connects frame selection, intent-aware layout planning, and background-conditioned text rendering through an explicit text–quadrilateral–color interface. The frame selector identifies a relevant and visually usable background. The planner first produces a natural-language rationale for title hierarchy, placement, and contrast, then predicts structured text regions and quantized colors. A diffusion renderer converts that plan into rasterized spatial and color conditions and uses localized OCR-feature supervision to strengthen character fidelity. We introduce Cover-100K, containing 99,781 aligned examples of recovered clean backgrounds, completed covers, strings, geometry, colors, and generated design rationales. On Cover-600 with supplied clean backgrounds, DecoupleGen improves text accuracy and distributional image quality over released editors, while human raters favor its planned layouts over PosterLlama. Component studies support the contributions of cover-domain supervision and structured rendering. Together, the dataset and structured interface provide a practical path from video content and design intent to readable covers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.