acceptodds
Under review as a conference paper at ICLR 2027

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Abstract

While coding agents are typically evaluated on end-to-end patch generation, their success heavily depends on an upstream context-acquisition stage: identifying relevant repository files. We introduce Agent Retrieval Bench, a file-level code retrieval benchmark designed for this workflow-driven context discovery problem. Unlike traditional semantic code search, relevance in Agent Retrieval Bench reflects what an agent needs next across four core tasks: code2test, comment2context, trace2code, and edit2ripple, alongside a selective retrieval suite with natural no-gold and counterfactual controls. The benchmark comprises 427 curated samples across 25 repositories over frozen base commits. Evaluating lexical, dense embedding, and RepoMap baselines reveals sharp task-dependent trade-offs and limited selective success on natural no-gold labels. Logged agent trajectories miss all gold files on 27–35% of samples. In a 45-sample, single-run pilot, retrieval-derived seeds achieve higher final File F1 with fewer post-seed calls and read tokens than random non-gold seeds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.