HxAgent: Iterative LLM Agent Planning for Generating Replayable Web Test Scripts from Natural Language
Abstract
Turning natural-language functionality descriptions into executable test scripts is a labor-intensive step of web UI testing. We present HxAgent, an iterative LLM-based planning agent that executes a described task in a browser and emits a replayable, locator-based test script that runs without any LLM at replay time. After every action, a dedicated state evaluator reassesses progress and can stop a diverging trajectory, and the next action is chosen only among executable elements, conditioned on (1) the current observation, (2) a short-term memory of prior state–action pairs, and (3) examples and textual rules accumulated in context from HxAgent’s own earlier successful and failed attempts; no model parameters are updated. HxAgent reaches 97.4% Exact-Match on MiniWoB++ (10.5% above a state-of-the-art WALT), 83.8% on 350 real-world tasks (13.4% above WALT), and 50.7% on OnlineMind2Web (13.5% above WALT). Counting only scripts that are generated and successfully replayed, it succeeds end-to-end on 78.9% of real-world and 38.2% of OnlineMind2Web tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.