WUIDGET: Web UI Generation with Small Models via Gap-Driven Post-Training
Abstract
Web UI generation requires a model to produce executable code, coherent visual design, and working interactions in the same artifact. Small models fail differently across these axes, while aggregate benchmark scores provide little guidance about which training signal should address which failure. We propose *WUIDGET*, a gap-driven post-training framework that distills evaluation failures into a structured root-cause taxonomy and uses it to guide data and reward design across three stages. Supervised fine-tuning learns broad structural and aesthetic patterns; Rejection Fine-Tuning with Distillation (RFT-D) minimally repairs the model's own rollouts; and reinforcement learning optimizes verifier-grounded visual and interactive rewards. On OpenDesign (840 prompts), the final Qwen3-8B checkpoint improves over its base from to static ( points) and from to agentic (a % relative gain in working interactions per artifact), and improves over the comparably sized AesCoder-4B by static. The static improvement over AesCoder remains consistent on WebDev Arena () and Design Arena (). Across Qwen3-8B and SmolLM-3B, SFT supplies the largest visual-quality gain, RL the largest interactive gain, and inserting RFT-D changes the static–agentic trade-off reached by downstream RL. Matched gap/no-gap controls show that SFT guidance improves the macro-average agentic score by (95% bootstrap interval ), while gap-selected RL queries produce a smaller, primarily static shift. These results support gap analysis not only as diagnosis, but as an organizing signal for specializing small models on multi-axis generation tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.