What Does a Failed Rollout Supervise? Separating Policy Failure from Recoverability in Tool Agents
Abstract
When a tool-using agent fails to complete a task, does that mean its earlier actions left no way to succeed? A failed continuation may reflect the limitations of the policy rather than an unrecoverable state. This distinction matters when sampled outcomes are used to label intermediate actions for training. We develop an executable audit that separates policy-relative success from recoverability within a specified continuation family and budget. Successful continuations justify positive recoverability labels, complete bounded non-reachability certificates justify negative recoverability labels, and unresolved cases remain unlabeled. Exhaustive replay over finite action inventories under fixed continuation bounds supplies exact audit labels in small deterministic environments. Our analysis distinguishes population identifiability from finite-sample certification and connects target mismatch to probability loss and representation-dependent ranking. Audits of released supervision data and six target builders show why sampled failure and native negative labels cannot uniformly be interpreted as irrecoverability. In a three-seed, four-arm Retail intervention with restricted candidates and matched training budgets, exact relabeling reduces held-out recoverability log loss by 0.150 relative to forced one-rollout labeling, while certificate-based abstention reduces it by 0.050 relative to matched deletion. Task-held-out monotone calibration of all arms, using exact labels from other tasks, narrows these gaps to 0.024 and 0.011 while preserving within-task rankings. Effects on action ranking and terminal success remain inconclusive. These findings motivate target-explicit supervision and separate evaluation of probability fit, action ranking, and terminal task success.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.