STEP: Agent Confidence Estimation from Execution Trajectories
Abstract
Agent confidence estimation (ACE) estimates how likely a Large Language Model (LLM) agent is to complete its task and is crucial for reliable agent deployment (e.g., web agents placing orders or clinical agents assisting diagnosis, etc.). However, most existing ACE methods remain difficult to deploy in practice due to three limitations: (1) Insufficient evidence. They derive confidence from token probabilities or self-reported certainty, which reflect how confident the model is in generating its text rather than whether its actions achieved the intended effect in the environment; (2) Biased information. They impose excessive hand-crafted inductive biases, resulting in the loss of raw agent information; and (3) Late estimation. They conduct estimation only after execution ends, offering limited support for early prediction. To address these limitations, we propose STEP (Stepwise Trajectory Encoding for Prediction), a framework that encodes the task, observation, and action text of each step to ground confidence in agent-environment interactions, applies temporal modeling to these raw step representations instead of relying on hand-crafted features, and employs multi-prefix supervision to support early prediction. Extensive experiments on ScienceWorld, WebShop, and InterCode Bash demonstrate that under within-task evaluation, STEP improves AUROC over the strongest baseline for each prediction setting by 19.2% on end-of-trajectory prediction and 11.6% on the more challenging early prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.