A Multidimensional View of LLM Uncertainty Quantification via Scenario Interventions
Abstract
Reliable uncertainty quantification for large language models should capture not only whether an answer changes, but whether it changes appropriately when the supporting evidence is altered. We propose Response–Evidence Coupling (REC), a multidimensional framework that estimates answer uncertainty from two coupled dimensions: scenario-conditioned response behavior and scenario-specific evidence support. Starting from an anchor answer, REC constructs evidence-preserving equivalent scenarios and counterfactual replacement scenarios that modify answer-bearing facts. It then evaluates whether the model follows the task-dependent expected response under each intervention and whether the generated answer is supported by the corresponding scenario. These signals are organized into anchor, equivalent-intervention, and counterfactual-intervention risks and aggregated through rank-normalized Fisher fusion. Across four QA benchmarks and four LLMs, REC achieves the best AUROC and AUARC in most settings. Ablation results further show that intervention-aware response behavior and scenario-specific evidence support provide complementary information, supporting a multidimensional view of LLM uncertainty.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.