Before You Rank Skill Memories, Measure What Your Benchmark Can Express
Abstract
Agent memory and skill systems are ranked by success rate on benchmarks whose capacity to express the gaps they report is never measured. We report that capacity for a controlled skill-form protocol, in which the same source experience is serialized four ways—trajectory trace, flat script, hierarchy, and a SKILL.md artifact—under a matched retriever, prompt, and injection budget. Four quantities govern what such a comparison can show: the error rate, the repeat-measurement floor (the rate at which per-task outcomes flip when one configuration is rerun), the discriminating fraction (the share of tasks whose outcome moves across arms for reasons a rerun does not explain), and whether the task slice represents the split. On an archived version of this protocol, which controlled variance by deterministic decoding, the floor is 14.7% of per-task outcomes on ALFWorld and 12.9% on WebShop, and only 3 of 50 WebShop tasks respond to skill injection at all. A rerun inflates the number of discordant pairs, never their direction, which motivates a practical reliability rule: a gap should exceed the control's own floor. On the published WebArena Shopping numbers of a recent budget-matched study of web-agent memory, none of its nine vanilla-versus-module gaps clears the vanilla actor's floor under any of its three executors. Rerun under a protocol that passes these checks (the full split, three replicates, a bank where 89 of 140 tasks retrieve a same-type source, length-matched fillers, two executor families), the serialization of the same experience, which bundles its actions, plan, and observations, decides whether it helps. A raw trajectory is indistinguishable from no injection on both executors and on every slice, while the hierarchical form of the same experience adds 15.7 and 19.3 points (p ≤ 0.001) and beats its own length-matched filler by 10.7 and 31.4 points. On one executor that gain needs a retrieved experience of the right task type; on the other it does not.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.