When Repairs Stop Being Recipes: Re-Grounding Repair Memory for Coding Agents
Abstract
Recent coding-agent systems store patches and debugging strategies from resolved issues as repair memory, then retrieve these records to guide later repairs and avoid repeated investigation. However, reusing a past fix can misdirect the next repair: a record may identify useful code while recommending an edit that no longer resolves the current issue. This repair applicability gap raises a question: how can an agent benefit from an earlier investigation when its fix no longer works? To study this question, we introduce RepairMem, a benchmark that pairs 99 current repair tasks with related historical repairs, covering both controlled faults and real repository issues. We evaluate five coding models on these tasks and find that giving them historical records unchanged usually reduces repair success compared with providing no memory. To make these records useful for the current task, we propose ReGround. Before the agent begins its repair, ReGround checks the historical patch against the current tests, uses the outcome to revise the historical repair advice, and supplies relevant code from the current repository to guide further debugging. This simple procedure requires at most one replay of the historical patch and no additional model calls for preparation. Across five models and 99 repair tasks, ReGround improves aggregate repair success by 16.4 and 27.7 percentage points over memory-free repair and direct historical reuse, respectively, while using 20.2% and 22.5% fewer model calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.