Harness²: Evaluating Coding-Agent Memory After a Reset
Abstract
A coding agent receives a stored code change to help repair its next program. The resulting recovery score is taken as a verdict on what the memory transfers. The harness can settle that verdict before a program reaches the verifier. On two held-out functions at a 16,384-token per-call cap, 19 of 24 episodes at extra-high reasoning effort emit no program; forced closure prevents an observed end-token failure mode on Qwen3.8-27B-FP8. We introduce Harness² (Harness of Harnesses), which starts from a failing program, excludes prior interactions and the target solution, and compares reference-derived banks with the task, seed, model, verifier, and execution limits held fixed. On 60 completed PACE-Bench task–seed units, a sibling reference change resolves 23 episodes versus 22 with no bank: 12 paired wins, 11 losses, and 37 ties, a difference of 1/60. A character-matched unrelated bank resolves 21, while direct patch application and whole-program transplant both fail on the tested cross-task transfers. Recovery is reported beside the output tokens consumed by the same evaluated episodes, excluding earlier executions that were replaced: the bank is evaluated through the harness that reads it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.