acceptodds
Under review as a conference paper at ICLR 2027

ContextGraph: Improving Coding Agents with Agent-Selected Lessons from Experience

Abstract

Coding agents repeat costly debugging mistakes even when a lesson that would avoid them was learned on an earlier task, often in another repository. We study what makes experience memory help such agents, using ContextGraph, an experience graph that distills successful and failed trajectories into strategies and warnings with the conditions under which they apply, lets the agent select them with its full context before planning and after errors, and adds every completed task back to memory. ContextGraph improves over no memory by 7.6, 13.7, and 11.3 points on SWE-bench Verified, SWE-ContextBench Related-Lite, and Datacurve DeepSWE, reaching 69.2%, 35.1%, and 72.1%, the highest rate among the evaluated memory methods on each. Its ablations show that two choices carry the gain: memory should store abstract lessons and let the agent select them for each task. On Related-Lite, raw trajectory excerpts selected by the agent and a fixed set of abstract lessons given to every task both stay at the no-memory level (20.5% and 20.8% versus 21.4%), and abstract lessons retrieved by embedding similarity leave ContextGraph only 2.9 of its 13.7 points. Updating memory after every task adds further gains, and lessons from other repositories alone carry more than half of the gain on Related-Lite and DeepSWE. The gains hold for strong agents: with Gemini 3.1 Pro, ContextGraph raises SWE-bench Verified pass@1 from 82.3% to 88.9%, above the 82.8% that five attempts reach without memory. We release the code and prompts so that others can build on these findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.