Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
Abstract
LLM agents operate in stateful environments, where a single erroneous step can waste limited interaction budget or cause irreversible effects before task failure becomes apparent. Reliable deployment therefore requires *step-level confidence estimation*: estimating, *before execution*, the probability that a proposed action will advance the task. Existing LLM confidence estimators are typically designed for static question answering under a fixed task context and evaluation criterion. For an agent, however, its action productivity depends on an environment transition that is observed only after execution. To address this challenge, we introduce Critic Experience Bank (CEB), a training-free framework that turns feedback from completed trajectories into reusable evidence for future confidence judgments. After each trajectory, an LLM assigns hindsight productivity pseudo-labels to individual actions and stores them with the critic's original pre-execution confidence, task context, action, and observed feedback. For a new action, by retrieving related productive and unproductive experiences, CEB grounds pre-execution confidence in feedback from completed trajectories to condition a fixed LLM critic. CEB thereby adapts over a task stream without parameter updates or ground-truth step labels at deployment. Across four agent benchmarks spanning offline and live web navigation, mobile GUI and shell tasks, and three critic backbones, CEB achieves the best or tied-best ECE, Brier score, and AUC in all twelve benchmark–backbone settings under rule-based step labels, reducing ECE by up to 53.8% relative to the strongest training-free baseline. Its confidence scores also improve downstream utility in selective execution and simulated task success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.