Budget-Aware Counterfactual Recovery for Long-Horizon AI Agents
Abstract
Long-horizon AI agents act through learned policies, tools, and external environments, but the action proposed by a base agent can be unreliable and runtime self-correction is not free. An external reliability controller can challenge a suspect action by restoring the current state, executing alternative recoveries, and comparing their realized continuations before committing. Because these counterfactual checks consume model, tool, and environment resources, they cannot be applied indiscriminately. We study when an agent action is worth challenging under a finite fork-execution budget, while tracking replay and critic costs separately, and learn state-conditioned recovery value for budget-aware runtime intervention. In a fault-injected tool-agent sandbox with noisy diagnostics, learned allocation raises held-out episode success from 9.3% without runtime forks and 14.1% with equal allocation to 17.0% at a fork budget of 12. The same learned allocation idea transfers to Taxi-v4 without latent-state access: at budget 24 it improves deadline-constrained success over direct execution by 13.9, 10.1, and 6.5 percentage points at action-error rates 0.10, 0.20, and 0.30 while using fewer forks than a hand-designed reserve-aware policy; the 0.40 condition is inconclusive and reserve-aware retains higher success overall. On a pre-locked 24-task ALFWorld OOD set with Gemini 3 Flash, direct execution succeeds on 21/24 tasks and selective recovery policies on 22/24, with fewer interventions than fixed forking but no population-level superiority claim. We then freeze an agent-specific learned challenge model on the development set and evaluate it on a second untouched 24-task OOD set with exact repeated base-policy prompts shared across methods. Direct, reserve-aware, and learned challenge each succeed on 23/24 tasks; learned challenge uses 1.50 forks per task versus 2.50 for reserve-aware, but does not rescue the sole direct failure and incurs 16.75 replay steps versus 14.92. Recovery can also harm otherwise successful trajectories, and fewer forks need not imply lower total runtime cost when restoration requires replay. The evidence supports counterfactual challenge as a bounded inference-time reliability mechanism rather than an unlimited local repair operation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.