Inert Priors: Auditing Relevance-Selected Memory Priors Against Their Own Measurement Floor
Abstract
LLM agent memory systems increasingly inject a handful of procedural priors—adapted from skills, past experiences or hand-written operations—into their retrieval planner, and choose those priors by relevance to the query. An earlier study reported that priors adapted from public agent skills lower judge accuracy relative to no priors, and called this a negative-transfer boundary. That priors-versus-no-priors comparison cannot separate three explanations: misaligned knowledge, an out-of-domain bank, or a lossy adapter. We first replace that comparison with one that can. It compares relevance-selected priors against priors drawn uniformly from the same bank, through the same adapter, at the same budget and injection point, so only the selection rule varies. On the earlier study's own outputs this contrast is larger and more stable than the one it reported. We then measure what that study never did: how far two executions of the identical configuration land apart. The earlier data contain such a pair, and its spread is the same size as every effect the study reported. Re-running the identified comparison on LongMemEval at with repeated independent executions per arm on three current backbones, we find the reported effect excluded on the one backbone with enough repeats to test it, and unsupported on two with fewer. On the backbone with the most repeats, a bootstrap over questions and executions bounds the relevance-versus-random difference within points and excludes the the earlier study reported. On two further backbones the difference is small and of opposite sign, though two runs per arm cannot exclude that size. Single runs of a selector temperature that interpolates between the two rules, an equal-budget raw prior format, an LLM re-ranker and a second benchmark scored without an LLM judge all land inside the same floor. The relevance ranking does degenerate—over a bank with nothing relevant in it, three of eighteen skills take half of all injected slots—but its cost is the relevance-versus-random contrast already bounded above. We report the audit, the floor and the prospective readings, and argue that results of this kind should be read against a measured floor, and a stated smallest effect of interest, before they are read at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.