All-in-Image: An Image-Native Generator for Graphic Design
Abstract
Graphic design generation combines requirements for content, layout, and typography with visual references whose appearance should be retained. Existing systems commonly encode these conditions through separate language, layout, and reference interfaces. Our observation is that different control objectives can share a unified semantic representation. Various requirements can all be expressed on a single spatial canvas, which describes the desired design without depicting its final appearance. Based on this observation, we introduce All-in-Image, a unified visual conditioning method that represents design conditions with two aligned images. On the semantic image, a global prompt, regional prompts, and text-rendering instructions are drawn together with their target regions, allowing semantic content and spatial extent to be represented in the same image. On the appearance image, subject and glyph references are placed at their target locations to provide appearance and glyph structure. A semantic encoder reads the written instructions and their spatial context, while a pixel encoder extracts features that preserve reference appearance. Their tokens jointly condition a single-stream diffusion transformer, allowing language, spatial regions, and appearance references to enter the generator through image space without a conventional text encoder or numerical bounding-box inputs. Experiments on the CreatiDesign benchmark show this formulation supports spatial control, text rendering, and subject fidelity jointly, ranking first on nine of ten metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.