Think or Call? Counterfactual Necessity Optimization for Software Agents
Abstract
Software agents must repeatedly decide whether to invoke a tool or continue reasoning with the evidence already available. Terminal task rewards provide limited credit for individual invocations: successful trajectories may contain redundant calls, while failed trajectories may contain essential ones. We introduce Counterfactual Necessity Optimization (), a method that learns tool routing by comparing the downstream value of these two choices. At selected decision states, restores the same agent and repository state into two isolated branches, executes a proposed tool call in one and a reasoning step in the other, and completes both under a shared frozen policy. Their return difference, accounting for task completion, protocol compliance, and optional-call cost, supplies a soft necessity target. A necessity head learns these targets jointly with the policy and guides routing at deployment without counterfactual rollouts. An obligation mask enforces required interactions during training and inference. We establish conditions for unbiased paired estimation and show that uniform error in the optimal tool-versus-reason value gap yields routing regret of at most over horizon . On SWE-bench Verified, improves resolution from 38.8% to 45.8% over outcome-only reinforcement learning while reducing optional calls per resolved issue by 39.1%. Further evaluation on SWE-Bench Pro shows consistent resolution gains. Routing diagnostics show fewer unnecessary invocations and skipped necessary calls, with protocol validity close to mask-only training. These results support invocation-level counterfactual credit for improving task completion and tool efficiency under explicit interaction constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.