Does Early Failure Prediction Actually Depend on Agent Execution?
Abstract
Early failure prediction aims to identify language model agent runs that are likely to fail. Yet a predictor can rank outcomes well without relying on the actions and feedback from the run being monitored. We test this distinction through task-matched process replacement. We keep the fitted predictor and the initial request fixed, then replace the visible process with one from another run of the same task and configuration. Replacements are chosen without using outcome labels. On 982 held-out-task runs from the public τ²-Bench archive, a tree predictor achieves an AUROC of 0.694 before the third action. Its mean AUROC is 0.692 when given another run's process from the same task, but falls to 0.652 when the process comes from a different task. Across three predictor settings and two early observation points, same-task replacement leaves mean AUROC close to or above the original value. Cross-task replacement lowers mean AUROC in all six comparisons. Same-task Brier score comparisons show no clear advantage for the original process. A separate analysis of 1,180 public SWE-Gym runs finds that adding process features improves a tree predictor, but an initial-only predictor achieves similar AUROC. These findings separate overall failure prediction from dependence on the current execution. Task-matched replacement provides a direct check of whether a predictor benefits from observing the particular run it is meant to monitor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.