acceptodds
Under review as a conference paper at ICLR 2027

StopLoss: Teaching Agents to Predict Failure and Stop Early

Abstract

Coding agents keep working on runs that will ultimately fail verification, and asking a frontier model to watch for this is expensive in its own right. We introduce StopLoss, a small external monitor that learns failure recognition from completed trajectories and their verifier outcomes. Final outcomes label earlier local windows, so a small fine-tuned language model can score the most recent tool steps in a single forward pass, and a threshold calibrated on successful runs turns that score into an abort decision without changing the executor. Trained on GPT-family SWE-rebench trajectories and evaluated on SWE-bench Verified and a SWE-bench Pro subset with GPT and Claude executors, StopLoss is compared against six general-purpose judges reading identical windows. Outcome supervision lifts failure prediction from near chance to a useful signal at the first check, at one to two orders of magnitude lower cost per check than any tested judge. This translates into savings: relative to completing every task, StopLoss lowers cost per retained solve by 4.5% on Verified while preserving 99% of successful runs, and by up to 31% on the costlier Pro subset, and in all four settings its cost per retained solve is below that of every judge policy. Ranking transfers less well to the unseen Claude executor, and a threshold calibrated on one benchmark does not carry its false-kill rate to another. Outcome-trained small monitors can thus make failure recognition cheap enough for early stopping to pay for itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.