How Much Can Successful Behavior Change Before Success Breaks?
Abstract
An LLM agent can complete the same task in different ways. Can one successful behavior be changed into another, one valid edit at a time, while preserving reliable task completion? Such flexibility matters when agents take detours or combine previously successful decisions. We introduce interventional success regions to study this question. Restartable checkpoints are linked by valid edits to their action histories and belong to the success region when the agent can reliably finish the task from them. We prove that success rates and specified local measurements can agree while successful connectivity differs. We develop a budgeted search procedure and test discovered paths with fresh continuation trials. In ALFWorld and τ-bench retail, adding connectivity information improves predictions of task success on unseen task families, reducing Brier loss by 3.37%. The largest gains occur when initial-condition and interface changes are combined. The average predictive benefit extends to models and task families unseen during predictor training, although gains vary across models. We propose a recovery training method that selects checkpoints using success-region structure and trains agents to predict recovery success. Across three models, the resulting procedure raises success under changed task conditions by 5.12 percentage points over a baseline matched for training difficulty. When both methods use the same checkpoint selection rule at test time, the gain remains 4.19 points [95% CI: 1.88, 6.46], supporting a benefit from training itself. These results show that success-region connectivity provides useful information for predicting task success and guiding recovery learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.