acceptodds
Under review as a conference paper at ICLR 2027

WebMetaBench: Interactive Black-Box Testing as Multimodal Meta-Evaluation for Web Development Agents

Abstract

Web development agents are increasingly capable of end-to-end application development beyond isolated code generation. Their expanding scope makes reliable evaluation critical and challenging, motivating systematic assessment of agentic judges. Existing evaluation relies mainly on structured page information or static screenshots, while many frontend requirements involve dynamic visual behaviors such as particles, animations, and interactive responses. Evaluating such behaviors requires a judge to actively test the application, acquire diagnostic video evidence, and reason over its dynamics. We introduce WebMetaBench, an interactive black-box benchmark for multimodal meta-evaluation of web development agents. It contains 120 single-defect and 19 multi-defect instances, with a gold-video setting that decouples active testing from video understanding and enables fine-grained evaluation of judge reliability. We further propose a three-stage failure paradigm consisting of requirement coverage, evidence acquisition, and video reasoning for fine-grained diagnosis of agent capabilities. Our experiments reveal substantial limitations in current models, with the best model achieving only 25.8% accuracy on single-defect instances and 26.9% defect recall in the multi-defect setting. WebMetaBench exposes critical limitations in both active testing and video understanding, providing a foundation for developing reliable judges for complex web development.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.