DigiDeck: Evaluation of Visual Web Agents in Synthetic Interactive Environments
Abstract
Evaluating multimodal web agents on the live web is challenging due to limited reproducibility, bot detection, and the potential for irreversible financial, privacy, or security consequences. Sandboxed environments mitigate these risks, but incur high development overhead and often fail to capture the diversity and complexity of real-world web interactions. To address this issue, we introduce DigiDeck, an automated framework for generating synthetic, interactive web environments for safe and reproducible evaluation of visual web agents. DigiDeck uses coding LLMs as web world models to predict action-conditioned state transitions in the form of browser-renderable environments, enabling interactive synthetic web trajectories without relying on live services. We demonstrate DigiDeck's utility by evaluating how agent performance varies across two dimensions: domain shifts and intent requirements. Using Online-Mind2Web as an anchor, we construct a benchmark of 1,080 tasks that systematically varies both intent (navigational, informational, and transactional) and domain distributions. Across seven agents, models show a consistent performance hierarchy across intents: performance decreases by 20.6 percentage points on average from navigational to informational tasks and by a further 17.8 points from informational to transactional tasks. On the distribution-shift task set, aggregate performance decreases by 5.2 percentage points, with substantial variation across models and a larger average informational-to-transactional gap.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.