acceptodds
Under review as a conference paper at ICLR 2027

When Does Continuation-Conditioned Supervision Improve Tool-Using Agents?

Abstract

Counterfactual supervision values an action by executing the remainder of a task with a chosen controller. Even an exact return can favor an action that is poorly suited to a different continuation. We study this effect in a controlled tool-use simulator, holding checkpoints, candidate actions, and the current-action risk rule fixed. We derive sufficient conditions under which continuation changes preserve the learner-visible optimal action, accounting for fitting error. Fixed-checkpoint diagnostics show that reference-conditioned selection is strictly suboptimal under a frozen deployment continuation at 16.8% of 800 checkpoints. We then evaluate newly fitted policies on held-out complete tasks, keeping training states and fitting procedures fixed across targets. Relative to reference supervision, fully deployment-conditioned targets yield mean paired utility gains of 0.076 for a cost-sensitive classifier and 0.099 for a multilayer perceptron. Both learner classes gain 3.0 success percentage points on average over three seeds, with improvements in utility and success in all six matched comparisons. These findings show that continuation choice can affect both the actions favored by supervision and the complete-task performance of policies learned from it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.