acceptodds
Under review as a conference paper at ICLR 2027

Action-Equivalent Memory Sketches for Web Agents

Abstract

Web agents powered by large language models often fail not because the underlying model is too weak, but because the memory that conditions the policy is semantically faithful yet action-irrelevant. We formalize this gap as an action-equivalent memory sketch: a compact history representation that preserves exactly the information capable of changing the next action, operationalized as an eight-field structured sketch (goal, sub_goal, completed, failed, constraints, entity_bindings, pending, evidence) updated at every step by a single LLM call. Three results bound its content: a soundness theorem for policy equivalence, an information-theoretic lower bound on sketch size, and a minimality theorem showing the eight fields are the unique action-complete cover. We evaluate on Mind2Web test_website across three runs. On qwen-turbo (10 task, 12 conditions) ours reaches Step-SR 0.0964; the B3 diagnostic, which removes constraints and failed, drops Step-SR to 0.0723, a 25% relative decrease consistent with the minimality theorem. On GLM-5.3-Flash (10 task, 8 per-field ablations) the direction reverses: B3 reaches 0.1233 versus ours 0.1096, with sub_goal the largest up-move and constraints the only down-move, opposite to what the theorem predicts. Suspecting a sample-size artifact, we ran a 50-task replication restricted to B3 and ours (438 step-judgments per condition, horizons 2 to 16). The direction returns to ours above B3 at 0.1347 versus 0.1187, a 13.5% step-level gain. The two-backbone picture is that the sketch works on both backbones once the sample is large enough, and 10-task direction flips are sample-size artifacts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.