Action-Equivalent Memory Sketches for Web Agents
Abstract
Web agents powered by large language models often fail not because the underlying model is too weak, but because the memory that conditions the policy is semantically faithful yet action-irrelevant. We formalize this gap as an action-equivalent memory sketch: a compact history representation that preserves exactly the information capable of changing the next action, operationalized as an eight-field structured sketch (goal, sub_goal, completed, failed, constraints, entity_bindings, pending, evidence) updated at every step by a single LLM call. Three results bound its content: a soundness theorem for policy equivalence, an information-theoretic lower bound on sketch size, and a minimality theorem showing the eight fields are the unique action-complete cover. We evaluate on Mind2Web test_website across three runs. On qwen-turbo (10 task, 12 conditions) ours reaches Step-SR 0.0964; the B3 diagnostic, which removes constraints and failed, drops Step-SR to 0.0723, a 25% relative decrease consistent with the minimality theorem. On GLM-5.3-Flash (10 task, 8 per-field ablations) the direction reverses: B3 reaches 0.1233 versus ours 0.1096, with sub_goal the largest up-move and constraints the only down-move, opposite to what the theorem predicts. Suspecting a sample-size artifact, we ran a 50-task replication restricted to B3 and ours (438 step-judgments per condition, horizons 2 to 16). The direction returns to ours above B3 at 0.1347 versus 0.1187, a 13.5% step-level gain. The two-backbone picture is that the sketch works on both backbones once the sample is large enough, and 10-task direction flips are sample-size artifacts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.