AppWright: Benchmarking End-To-End Application Construction
Abstract
Coding agents are increasingly deployed as products that build complete applications from natural-language requirements, yet existing benchmarks evaluate function-level correctness, repository-level issue fixing, or single-page web generation judged by human preference. None measures the end-to-end ability to deliver a full, runnable, and well-designed application across the forms industrial users request. We present APPWRIGHT, a benchmark abstracted, under a desensitization and deduplication pipeline, from the request traffic of a production no-code application-generation service, covering five categories: web apps, tools, games, mini programs, and mobile apps. APPWRIGHT provides (i) 184 tasks from anonymized real user requirements, each shipping the original query and a structured Product Requirements Document (PRD) with an acceptance rubric and a difficulty label (D1–D4); (ii) a reproducible construction-and-execution harness reimplementing the Code Arena specification; and (iii) a reward agent scoring each delivered application on two orthogonal axes—usability, verified by actively operating the deployed app, and aesthetics, a versioned nine-item rubric distilled from production design reviews. To our knowledge, APPWRIGHT is the first benchmark to jointly evaluate usability and aesthetics under structured, reproducible criteria. Across nine frontier models the two axes disagree: the top model differs by metric, difficulty rather than application category decides the ranking, and aesthetics discriminates only as a gate—so a single aggregate score would hide the differences that separate systems. We release the task suite, harness, and reward agent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.