Choosing the Task Order Returns More Than Governing What Gets Written
Abstract
Self-evolving LLM agents improve by writing their own experience into a memory bank and reusing it on later tasks. Prior work has focused on what the bank should keep: which experiences to admit, distil or prune. We show that a variable one step earlier can matter more: the order in which the same tasks arrive. On execution-graded code benchmarks, permuting a training stream while holding its content fixed moves held-out success by 4-11 points, and the spread shrinks to evaluation noise once memory is disabled, locating the effect in what the agent writes. Order changes which experiences exist, and filtering them afterwards does not recover the gain: at identical cost per deployed run, choosing the order adds 1.71 points on average, while write-time governance and seven further admission and retrieval rules yield no significant gain. We turn this finding into the ordering lever, a pilot protocol that briefly tries each candidate order on the target task pool and commits to the winner. It raises all four write policies on HumanEval+ by 2.3-3.2 points, including Reflexion, ExpeL and Agent Workflow Memory, recovers over 90% of the best candidate's gain on 34 freshly drawn task pools, and comes with a reliability score, computed from the pilot alone, that tracks its payoff. The best order is specific to each writer: installing one writer's winning order on another usually hurts, while that writer's own pilot helps. For agents that learn from their own experience, arrival order is a design choice to be measured, not copied.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.