acceptodds
Under review as a conference paper at ICLR 2027

VEA-Aware Uncertainty for LLM Chain of Actions

Abstract

For the prediction of a binary outcome, Variability–Epistemic–Aleatoric (VEA) uncertainty quantification separates the error of a single model-reported probability into three components: the variability (V) of the prediction across repeated runs of the model, the epistemic gap (E) between the average prediction and the true conditional probability of the outcome, and the aleatoric uncertainty (A) of the outcome itself, which no predictor can impact. This decomposition presents an important distinction that calls for different strategies to improve model performance: averaging over repeated runs of the model of interest reduces variability, and better modeling choices and training procedures reduce the epistemic gap while reporting the intrinsic aleatoric error. We propose an estimation procedure of VEA components in the setting of predictions through a chain of actions, namely a single structured generation mechanism that emits an ordered sequence of intermediate judgments before its final action. First, when the true conditional probability is unobserved, we estimate the epistemic gap (E) against a disclosed proxy model, and we provide an aggregate diagnostic that checks the adequacy of that proxy using the observed outcomes. Second, we further decompose the variability (V) into an exact sum of individual action-specific contributions so that action-level instability can be identified and targeted. We apply our VEA framework to evaluate the replicability of published scientific findings, an important but resource-intensive task in practice, using three core benchmark replication datasets (RPP, CB and SSRP). Across runs per study, the VEA decomposition shows that the overall prediction error is dominated by the epistemic component (E) and is concentrated in a small number of specific actions along the chain. This implies that increasing the reasoning effort of the LLM under evaluation improves prediction accuracy, whereas lowering the sampling temperature does not, because the already small variability (V) has no room for further shrinkage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.