acceptodds
Under review as a conference paper at ICLR 2027

Code Understanding is a Bottleneck for Coding Agents

Abstract

Repository benchmarks (e.g., SWE-bench) for coding agents often assume that edit volume can predict task difficulty, but their limited control over code and task types makes it hard to discern which abilities truly drive agent errors. We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluation. Our framework designs tasks from scratch as call graph transformations and scales their size over four axes: function traversal, search, runtime analysis, and instruction following. We evaluate eight LLMs and six coding agents on 6,840 CABRA tasks to show: 1) LLM accuracy falls as task size grows, but agents stay near-perfect by offloading work to tools (e.g., grep); 2) Larger CABRA tasks elicit more tool calls for reading and analysis, while on SWE-bench Verified these tool call counts predict agents' task accuracy better than edit volume—suggesting task difficulty for agents can lie in *understanding* code to edit, not solely in making edits; 3) Extending CABRA to an intense understanding task where models analyze divergent logic across two classes backs this hypothesis, as agent accuracy finally falls. Overall, we argue for revisiting synthetic evaluation protocols like CABRA to unmask LLM weaknesses trivialized by tools (e.g., needle-in-a-haystack) and abilities beyond editing (e.g., understanding) that coding agents still find difficult—complementing repository-based tasks with controlled diagnosis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.