acceptodds
Under review as a conference paper at ICLR 2027

Does Early Failure Prediction Actually Depend on Agent Execution?

Abstract

Early failure prediction aims to identify language model agent runs that are likely to fail. Yet a predictor can rank outcomes well without relying on the actions and feedback from the run being monitored. We test this distinction through task-matched process replacement. We keep the fitted predictor and the initial request fixed, then replace the visible process with one from another run of the same task and configuration. Replacements are chosen without using outcome labels. On 982 held-out-task runs from the public τ²-Bench archive, a tree predictor achieves an AUROC of 0.694 before the third action. Its mean AUROC is 0.692 when given another run's process from the same task, but falls to 0.652 when the process comes from a different task. Across three predictor settings and two early observation points, same-task replacement leaves mean AUROC close to or above the original value. Cross-task replacement lowers mean AUROC in all six comparisons. Same-task Brier score comparisons show no clear advantage for the original process. A separate analysis of 1,180 public SWE-Gym runs finds that adding process features improves a tree predictor, but an initial-only predictor achieves similar AUROC. These findings separate overall failure prediction from dependence on the current execution. Task-matched replacement provides a direct check of whether a predictor benefits from observing the particular run it is meant to monitor.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.