acceptodds
Under review as a conference paper at ICLR 2027

A statistical measure of trust in generative AI

Abstract

Trustworthiness is a crucial issue for generative AI, as modern models can produce convincing and confident responses even when their outputs are incomplete, unsupported, or wrong. The frequent absence of ground-truth knowledge or labels complicates verification of generative AI outputs. A further complication is the opacity of training data: in practice, we often cannot distinguish between a model well-trained on low-quality data and a model trained on high-quality data whose internal states nonetheless produce atypical outputs. Trustworthiness can be approached through the lenses of uncertainty, consistency, and confidence – concepts frequently conflated despite their distinct operational meanings. We argue that clearly differentiating them and advancing confidence-based reliability estimation is essential for progress. We propose a trustworthiness framework grounded in model-internal confidence calibrated against an external expert reference. We estimate **per-token trustworthiness** of a student generative model as the statistical typicality of each generated token relative to an expert model's distribution, where the expert is used as a distributional reference rather than as a source of factual verification labels. To this end, we introduce an **AI ontology calibrator** that leverages the student model's internal latent features and token-level output statistics to compute expert-reference acceptance scores. Theoretically, these scores are supported by statistically grounded bounds on threshold-defined acceptance and rejection probabilities under an expert-reference calibration criterion; they should not be interpreted as probabilities of factual correctness. Combined with traditional token-level confidence, entropy, and perplexity statistics, calibrator scores yield a compact **set of 15 response-level indicators** for correctness classification. Evaluated across BF program generation, mathematical problem solving, and medical question answering, combined calibrator and traditional indicators provide complementary reliability signal which could be used to build detectors of untrustworthy outputs and hallucinations. Remarkably, our experiments show that these detectors are transferable across domains, although transferability is not always asymmetric.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.