Learning to Draw Interfaces: What Makes Synthetic UI Data Work for Text-to-Image Models
Abstract
Generating interface screenshots requires text-to-image models to reproduce both the specified text and the arrangement of components. In this paper, we propose ScreenForge, a data synthesis system for studying how training data shapes these capabilities. The system renders application templates and retains source documents and component trees, allowing different captions to be derived for the same image. We use ScreenForge to compare caption recipes, application coverage, and UI-to-background mixtures in single-run continued-training experiments, each starting from the same 2B checkpoint and using 10 million examples. These comparisons show that caption choice depends on the inference prompt format and that broader application coverage improves most metrics at a fixed UI sample count. Increasing the UI share does not improve all metrics: moderate shares favor text and layout adherence, while larger shares yield modestly higher visual quality. On a benchmark of 1,300 structured screen specifications, a 50% UI mixture raises required-text recall from 0.465 to 0.813 and VLM-rated layout adherence from 3.24 to 3.84 compared with background-only continued training. We will release the rendered training images and captions, data manifests, and all 1,300 benchmark reference screens, together with the evaluation code, to support further research on UI generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.