acceptodds
Under review as a conference paper at ICLR 2027

CounterGraft: Counterfactual Execution for Agent Self-Improvement

Abstract

A higher-scoring agent can still make worse behavioral choices: stronger downstream execution can conceal unnecessary or harmful changes. We introduce **CounterGraft**, a method for improving executable agent policies and prompts through counterfactual feedback. Its core operation runs the successor from the incumbent's selected choice and execution state, then compares the outcome with execution from the successor's own handoff. Holding the successor policy fixed reveals whether the replacement improves declared utility, supplying revision feedback unavailable from paired endpoints alone. CounterGraft uses this feedback to repair programs, then independently audits the selected program before release. On AgentDojo, it improves held-out task success by up to 32.9 percentage points over incumbent agents and by 1.3–1.9 points over the strongest adapted baseline across three executors; it also achieves the highest success among the compared methods on ALFWorld. In a matched AgentDojo experiment, exposing counterfactual feedback reduces harmful replacements by 70% while increasing task success, with the proposer, executor, and proposal budget fixed. Repairs recover required actions and improve replacement quality at an unchanged departure count. The release audit controls false certification on the bound workload under adaptive stopping, reduces replay beyond deterministic checks, and enables certified successors to become subsequent incumbents in prospective runs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.