Metrics-in-the-Loop: Cheap, Timely Course Corrections Improve Coding Agents
Abstract
In spite of their impressive recent progress, coding agents still accumulate overly complex functions and untested code over long trajectories. Nothing flags these weaknesses as they form: the only signal is a failing test, and by then they likely have compounded. We introduce metrics-in-the-loop, an inference-time intervention that gives agents in-flight feedback: cheap, automatically computed code metrics injected into the agent's context during the trajectory. After each code-changing agent turn, we insert short feedback messages naming changed functions whose cyclomatic complexity exceeds a limit (complexity feedback) or changed lines that are skipped by tests (coverage feedback), giving the agent the opportunity to refactor code or strengthen tests before continuing. On two Python benchmarks, complexity feedback consistently leads agents to reduce the concentration of codebase complexity. It does so with no detrimental impact on functional correctness. Generic reminder messages, on the other hand, have almost no effect. Our experiments investigate a practical form of agent “re-railing”, where timely, actionable measurements can effectively redirect work before local weaknesses become embedded, suggesting practical paths for coding agent to work effectively on long-term-maintainable codebases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.