acceptodds
Under review as a conference paper at ICLR 2027

Learning to Draw Interfaces: What Makes Synthetic UI Data Work for Text-to-Image Models

Abstract

Generating interface screenshots requires text-to-image models to reproduce both the specified text and the arrangement of components. In this paper, we propose ScreenForge, a data synthesis system for studying how training data shapes these capabilities. The system renders application templates and retains source documents and component trees, allowing different captions to be derived for the same image. We use ScreenForge to compare caption recipes, application coverage, and UI-to-background mixtures in single-run continued-training experiments, each starting from the same 2B checkpoint and using 10 million examples. These comparisons show that caption choice depends on the inference prompt format and that broader application coverage improves most metrics at a fixed UI sample count. Increasing the UI share does not improve all metrics: moderate shares favor text and layout adherence, while larger shares yield modestly higher visual quality. On a benchmark of 1,300 structured screen specifications, a 50% UI mixture raises required-text recall from 0.465 to 0.813 and VLM-rated layout adherence from 3.24 to 3.84 compared with background-only continued training. We will release the rendered training images and captions, data manifests, and all 1,300 benchmark reference screens, together with the evaluation code, to support further research on UI generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.