Live-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents
Abstract
Training a browser agent for deployment requires assigning credit from a sparse terminal outcome to the few critical steps that determined it. Existing pipelines supervise on successful demonstrations and optimize online with group-relative objectives that assign a uniform trajectory-level advantage to every step. Methods resolving credit below the trajectory level either intervene on the environment or match steps via recurring observations, neither of which is reliable in live browsing sessions. We propose a two-stage training framework to address this. First, Recovery- and UI-specialized Curriculum Supervised Fine-Tuning (RUIC-SFT) treats recovery data as state-conditioned supervision, introducing it only after the conditioning state becomes separable; this significantly improves erroneous-step recovery while reducing redundant actions per task. Second, Divergence-Attributed Online Group Relative Policy Optimization (DAO-GRPO) derives unbiased step-level advantages from prefix-conditioned leave-one-out baselines within each rollout group, requiring no environment resets or per-step model calls. Building on these stages, we develop the Live-Browser-Agent family (4B, 9B, and 27B). Our 27B model attains the highest average Pass@1 among open-weight models on three online browser benchmarks (82.3%), on par with Kimi-k3 and Qwen3.8-Max despite using roughly 90× fewer parameters, and improves its Qwen3.5-27B initialization on three general agentic benchmarks, indicating that browser-centric training does not erode general tool use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.