ContextForge: Learning What to Retrieve From Diverse Validated Patches
Abstract
Search subagents can reduce the expensive repository performed by coding agents, yet they are commonly trained and evaluated using overlap with locations modified in a single solution patch. This proxy omits useful context in unmodified callers, interfaces, tests, and invariants. It also overlooks that there can be many alternative test-valid repairs that depend on different repository evidence. To address these limitations, we introduce a novel reward function for context retrieval that scores edit locations, supporting code, and explanations using patch-conditioned information gains averaged over multiple test-validated repairs. We implement this reward in ContextForge, a training framework that constructs diverse repair sets through offline generation and test execution. Using GRPO, we train a Qwen3.5-4B search agent and evaluate it with OpenHands and Mini-SWE-Agent on SWE-bench Pro, Verified, and Multilingual. Across all six benchmark-scaffold settings, ContextForge achieves the lowest inference cost and best or tied Pass@1 among evaluated context retrieval agents, and reduces end-to-end cost by up to 29.1% compared to the no-subagent baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.