TURF: Outcome-Grounded Uncertainty Quantification for LLM Agents
Abstract
LLM agents execute multi-step workflows with consequences beyond their final responses, making uncertainty quantification important for identifying tasks that warrant verification or human oversight. We introduce TURF, a task-training-free method that ranks task-level failure risk by combining smoothed dispersion of observable outcomes with frozen-NLI spectral diversity of terminal responses through Gaussianized-rank fusion, without fitting correctness-supervised predictors or fusion weights. TURF-Fixed uses a prescribed trajectory budget, while TURF-Adaptive selects a prefix using Bayesian evidence of outcome and semantic recurrence relative to the first execution. We evaluate both variants against 15 baselines on six benchmarks, three backbones, and three seeds, across supported trajectory caps from 2 to 16. Equally averaging across datasets, backbones, seeds, and budgets of 4, 6, 8, 12, and 16 trajectories, TURF-Fixed improves AUROC/AUPR by 5.19/2.99 percentage points over SGD, the strongest baseline with complete results on this grid. TURF-Adaptive retains gains of 4.68/2.60 points while using 20.90% fewer trajectories for scoring than TURF-Fixed, with corresponding quality losses of 0.52/0.39 points. The results support outcome-grounded, two-view uncertainty estimation across multiple sampling budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.