acceptodds
Under review as a conference paper at ICLR 2027

Before You Rank Skill Memories, Measure What Your Benchmark Can Express

Abstract

Agent memory and skill systems are ranked by success rate on benchmarks whose capacity to express the gaps they report is never measured. We report that capacity for a controlled skill-form protocol, in which the same source experience is serialized four ways—trajectory trace, flat script, hierarchy, and a SKILL.md artifact—under a matched retriever, prompt, and injection budget. Four quantities govern what such a comparison can show: the error rate, the repeat-measurement floor (the rate at which per-task outcomes flip when one configuration is rerun), the discriminating fraction (the share of tasks whose outcome moves across arms for reasons a rerun does not explain), and whether the task slice represents the split. On an archived version of this protocol, which controlled variance by deterministic decoding, the floor is 14.7% of per-task outcomes on ALFWorld and 12.9% on WebShop, and only 3 of 50 WebShop tasks respond to skill injection at all. A rerun inflates the number of discordant pairs, never their direction, which motivates a practical reliability rule: a gap should exceed the control's own floor. On the published WebArena Shopping numbers of a recent budget-matched study of web-agent memory, none of its nine vanilla-versus-module gaps clears the vanilla actor's floor under any of its three executors. Rerun under a protocol that passes these checks (the full split, three replicates, a bank where 89 of 140 tasks retrieve a same-type source, length-matched fillers, two executor families), the serialization of the same experience, which bundles its actions, plan, and observations, decides whether it helps. A raw trajectory is indistinguishable from no injection on both executors and on every slice, while the hierarchical form of the same experience adds 15.7 and 19.3 points (p ≤ 0.001) and beats its own length-matched filler by 10.7 and 31.4 points. On one executor that gain needs a retrieved experience of the right task type; on the other it does not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.