VETRA: Composable Evaluator Control for Repository Repair
Abstract
We treat a repository agent's last two evaluator calls as a transferable control problem: choosing a targeted test or a regression suite determines both what the agent can learn and whether it can validate the patch it returns. Verification and Evaluator-Control Transfer for Repository Agents (VETRA) makes that choice reusable: it stores Observe, Diagnose, Repair, Verify, and Budget as executable fields, routes on the latest failure, and deterministically composes compatible test order, coverage, call reserves, and stopping conditions. EvalRepair-320 supplies 320 tasks from 160 unseen repositories and a complete \(2\times2\times2\) intervention over transfer object, router evidence, and offline refinement while fixing the backend, online actions, bank capacity, and evaluator opportunities. At \(B=8\), VETRA achieves 84.7% success with 2 median evaluator calls, compared with 72.4% and 3 for the initialized typed router and 46.0% and 6 for bundled Meta-Policy Reflexion. Across the grid, typed-field system cells average 17.3 success points above bundled cells, failure-conditioned routing adds 9.9, and refinement adds 9.7. Verify+Budget leads all four prespecified equal-cardinality field pairs at both tested budgets. Hidden tests retain 79.5% success, blinded issue-satisfaction review yields 74.1% adjusted success, and frozen-bank gains persist on SWE-bench Lite and Verified-240. The result identifies evaluator policy—especially verification order and call reservation—as transferable state for repository agents operating with scarce execution feedback.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.