acceptodds
Under review as a conference paper at ICLR 2027

w1: Visually Grounded Code Improves Screenshot-to-HTML Generation

Abstract

This paper introduces w1, a very competitive vision-language model (VLM) for converting webpage screenshots into HTML code. We find this task bottlenecked by the quality of supervision targets: code of real webpages contains many tokens, such as invisible elements, unused styles, redundant nesting, and media source strings, which are not visible from the webpage screenshot. It is not possible for a VLM to predict such code from the screenshot. To address this problem, we propose a data pipeline that removes the invisible code while preserving the rendering results. Specifically, our pipeline parses the HTML, matches each DOM element with its computed CSS, removes code without visible effect, canonicalizes render-equivalent structure, and replaces media sources with placeholders. The result is the WebGround-3M dataset. It has 3M screenshot-HTML pairs, where the code grounds in the screenshot. In this dataset, supervision targets are 30% shorter on average while rendering very similar to the original screenshots. On WebGround-3M, fine-tuning a VLM demonstrates an effective improving trend. Moreover, from WebGround-3M, we curate 200k synthetic preference pairs targeting at syntactic repetition and semantic perturbation; direct preference optimization (DPO) using these preference pairs adds further performance gains. Across six public benchmarks, w1 achieves highly competitive code generation performance. In particular, on benchmarks evaluated using GPT scores and the Design2Code metric, w1-9B outperforms zero-shot GPT\-5 and Gemini-3.5-flash, and models trained on WebSight-v1/v2, VinciCoder, MCD, WebCode2M, and Web2Code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.