acceptodds
Under review as a conference paper at ICLR 2027

The Expression Gap: Probing What Language Models Know but Don't Say

Abstract

Instruction-tuned language models represent their own epistemic state—whether a claim is uncertain, self-contradictory, or unsupported by a provided document—more reliably inside the network than they reveal in generation. We call the difference between this internal representation and the model's output the expression gap, define it at a deployed threshold as probe fire-rate minus generation hedge-rate, and measure it across four open-weights families (Qwen2.5-7B, Mistral-7B-v0.3, OLMo-2-7B, Llama-3.1-8B). One recipe—a linear probe on the layer-17 last-input-token hidden state, fit per model and task—reads this state at AUC 0.97–1.00 on paired uncertainty and retrieval faithfulness across all four families, beats every input-only baseline (decisively on sycophancy prompts whose wording is identical across conditions), and transfers zero-shot across tasks within a model (off-diagonal mean AUC 0.96 on Qwen, 0.89 on Mistral). The gap it exposes is large and model-ordered (regex-scored 0.14–0.49; OLMo largest, Llama near zero under semantic judging) and instruction-gated: under a permissive prompt, retrieval unfaithfulness spans a 0.00–0.78 gap across families. Gating a sparse flag-token logit bias on the probe recovers expression with zero added false positives on the evaluation set (held-out gate exposure 1.7–6.3%); a blinded semantic audit of every stored generation (κ=0.94) confirms the lift on two families, exposes a third as surface-form conversion, and finds a null on the fourth. The obvious alternative—a one-line hedging instruction—lifts expression on every family but always at commensurate false-positive cost (+12 to +33pp), because a prompt cannot condition on what the model internally knows. A gated-actuator factorial locates the mechanism: the zero-FP property belongs to the probe gate, which every actuator behind it (logit bias, ITI, CAA) inherits. We release datasets, code, and the full judging record.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.