acceptodds
Under review as a conference paper at ICLR 2027

Evidence States Survive Where Answers Fail: Decodability and Causal Access in Language Models

Abstract

A record can support a claim, its explicit negation, both, or neither. Preserving these distinctions is essential for evidence reporting. We separate evidence-state recoverability, native reporting, and causal access using entity-role swaps that preserve the exact word multiset across all four support states. Across three instruction-tuned model families, the output contrast between absent and dual support is attenuated relative to the contrast between opposing one-sided support. On held-out signed records, same-query linear decoders distinguish absent from dual support with 87.5–100% accuracy, compared with 50.0–75.0% for native yes/no readout. This advantage persists over scalar decoders trained on identical examples and state labels. Residual interventions reveal task-conditioned causal access: a within-query state intervention in Mistral exceeds an equal-norm random control by 1.79 log-odds units. Averaged over both complementary questions, transferred directions underperform random controls on signed records. On unseen source- and time-scoped records, Mistral retains strong hidden-state decoding, while Qwen and Olmo show weaker transfer. Together, these findings establish a recoverability–reporting gap and identify question- and model-dependent boundaries on its causal portability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.