acceptodds
Under review as a conference paper at ICLR 2027

Attention as a Markov Random Field

Abstract

Reliably detecting when large language models (LLMs) hallucinate remains an open challenge. Existing uncertainty estimators either require costly multi-sample generation or supervised calibration data, limiting their practical deployment. We introduce the Attention Bethe Score (ABS), a zero-shot, single-pass uncertainty metric grounded in variational inference over probabilistic graphical models. Our key insight is to reinterpret the attention mechanism of a frozen Transformer as the structure of a Markov Random Field: token positions become nodes, attention weights define edges with attention-modulated Potts couplings, and hidden-state projections yield unary potentials. Running loopy belief propagation on this graph, we compute the Bethe free energy as a principled measure of internal representational frustration-the degree to which strongly attended positions hold inconsistent beliefs. We prove that, under a contraction condition satisfied in typical operating regimes, the Bethe free energy monotonically tracks projected frustration, and we validate this relationship empirically (Spearman p > 0.85). Across eight benchmarks spanning open-domain QA, fact verification, hallucination detection, and summarization, ABS achieves aggregate AUROC of 0.788 (LLaMA-3.1-8B) and 0.806 (Qwen-2.5- 7B), surpassing all baselines including multi-sample Semantic Entropy and supervised Internal Probes-while adding fewer than four milliseconds and less than three percent computational overhead to a standard forward pass.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.