acceptodds
Under review as a conference paper at ICLR 2027

ABPQ: Automated Construction of Challenging UI-to-Code Benchmarks via Round-Trip Consistency

Abstract

Existing UI-to-Code benchmarks face a tradeoff between evaluation comprehensiveness and construction scalability. Benchmarks containing paired UI-HTML data support fine grained objective evaluation but are costly to construct and are becoming increasingly saturated, whereas benchmarks containing only UI screenshots are easier to scale but typically rely on LLM-as-Judge. To address this issue, we propose PQ Score, which identifies samples that are challenging for models without requiring reference code by measuring information loss during the “screenshot → code → rendered screenshot → code” round trip process. Based on PQ Score, we further propose the ABPQ benchmark construction pipeline with minimal reliance on human effort and use it to construct ABPQ-FULL, containing 500 samples, and its challenging subset ABPQ-HARD, containing 200 samples, from 4,771,476 real UI screenshots. The final samples retain their corresponding HTML, supporting multiple automatic metrics, LLM-as-Judge evaluation, and human evaluation. Experiments on ten representative proprietary and open source models show that ABPQ is more challenging than representative existing benchmarks on nearly all metrics and that current advanced models remain far from saturation. In addition, PQ Score exhibits strong positive correlations with other automatic metrics, validating its effectiveness for difficult sample filtering. ABPQ provides a new approach to constructing scalable, continually updated UI-to-Code benchmarks with minimal reliance on human effort.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.