Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
Abstract
The web is complex, open-ended, and constantly changing, making it challenging to scale training environments for visual web agents. Existing data collection efforts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for reinforcement learning (RL), failing to capture the diversity of the web. Training directly on live websites could alleviate this, but it is slow, brittle, irreproducible, and can have unintended side effects on real websites. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. At its core, Weblica uses coding agents to synthesize interactive websites grounded in real-world websites and core web navigation skills, allowing us to scale RL training to thousands of diverse environments served locally. To assess how well synthetic websites can substitute for real ones, we further build a proxy for live-web training by caching real websites at the HTTP level, capturing and replaying stable visual states while preserving interactive behavior. Both synthetic and cached real environments yield substantial improvements over the base model, with purely synthetic RL achieving the best overall performance. Our best model, Weblica-8B, outperforms open-weight baselines of similar size across multiple web navigation benchmarks while using fewer inference steps, scales favorably with additional test-time compute, and is competitive with API models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.