acceptodds
Under review as a conference paper at ICLR 2027

CARVE: Replay-Verified Gradient Admission for Offline Agent Controllers

Abstract

Offline agent training often converts every logged operator decision into actor supervision even though replayability and behavior support vary across a trace. We introduce CARVE, Counterfactual Admission with Replay-Verified Evidence, a gradient-admission layer that tests joint operator–argument actions before objective construction and authorizes one of three paths: replayable supported decisions update actor and critic through counterfactual attribution, supported non-replayable decisions update only the critic, and unsupported decisions abstain. This tri-state contract retains supported value-learning signal while protecting actor gradients from replay-invalid and unsupported targets. With Llama-3.1-70B-Instruct and matched data, tools, objective, seeds, and token caps, CARVE raises Code/SWE success from \(60.9\pm1.2\) to \(63.8\pm1.1\), a \(+2.9\pm0.5\)-point paired gain with \(5/5\) positive seeds and task-clustered \(95%\) CI \([+1.5,+4.2]\); it also exceeds full-action BC+AWR by \(+1.4\pm0.4\) points with \(5/5\) positive seeds. An equal-footing admission-by-adapter matrix isolates routing: replay-and-support admission adds \(+1.3\) points with adaptation disabled and \(+1.5\) with it enabled, while the complete configuration exceeds the strongest adapter-matched simple rule by \(+0.8\pm0.4\) points. Region interventions select attribution in the replayable-supported slice and critic-only learning in the non-replayable-supported slice. Task-clustered gains remain positive across software, SQL, planning, and web workloads, and the same contract improves a distinct planner–executor controller by \(+2.5\) points on Code/SWE and \(+2.2\) on held-out Web. CARVE makes trace reliability an explicit, portable decision about gradient eligibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.