acceptodds
Under review as a conference paper at ICLR 2027

TURF: Outcome-Grounded Uncertainty Quantification for LLM Agents

Abstract

LLM agents execute multi-step workflows with consequences beyond their final responses, making uncertainty quantification important for identifying tasks that warrant verification or human oversight. We introduce TURF, a task-training-free method that ranks task-level failure risk by combining smoothed dispersion of observable outcomes with frozen-NLI spectral diversity of terminal responses through Gaussianized-rank fusion, without fitting correctness-supervised predictors or fusion weights. TURF-Fixed uses a prescribed trajectory budget, while TURF-Adaptive selects a prefix using Bayesian evidence of outcome and semantic recurrence relative to the first execution. We evaluate both variants against 15 baselines on six benchmarks, three backbones, and three seeds, across supported trajectory caps from 2 to 16. Equally averaging across datasets, backbones, seeds, and budgets of 4, 6, 8, 12, and 16 trajectories, TURF-Fixed improves AUROC/AUPR by 5.19/2.99 percentage points over SGD, the strongest baseline with complete results on this grid. TURF-Adaptive retains gains of 4.68/2.60 points while using 20.90% fewer trajectories for scoring than TURF-Fixed, with corresponding quality losses of 0.52/0.39 points. The results support outcome-grounded, two-view uncertainty estimation across multiple sampling budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.