Belief Elicitation Is an Intervention: Measuring and Manufacturing Belief–Action Consistency in Language Models
Abstract
Evaluating whether language models act consistently with their stated beliefs is central to assessing calibration, risk preferences, and safety alignment. Existing benchmarks typically measure this by asking a model for a subjective probability and observing its subsequent choice within the same conversation. However, treating this same-session agreement as evidence of behavioral predictability implicitly relies on an untested assumption: Measurement Non-Reactivity (MNR), which assumes that the act of asking a belief does not alter the subsequent action. We formalize this problem by separating the computed metric, same-session consistency (), from the true target, cross-context predictive validity (). We prove that fails to identify : under standard deterministic evaluation, the bounds on span the entire unit interval , rendering a perfect consistency score completely uninformative. To resolve this flaw, we propose the Shadow Fork, an evaluation protocol that pairs each elicited session with an unprompted action session to identify without assuming MNR, costing only one additional query per scenario. Across five frontier models, we show that standard protocols substantially inflate consistency due to reactivity: by in a calibrated delegation study and by on SimpleQA abstention. Crucially, interventional experiments injecting synthetic probabilities reveal that models mechanically conform to whatever number appears in their context ( decision flip rate). Therefore, observed belief–action consistency is largely an artifact manufactured by the evaluation prompt rather than evidence of genuine model introspection, highlighting the critical need for non-reactive benchmark design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.