acceptodds
Under review as a conference paper at ICLR 2027

Knowledge After Collapse: Context and Wording Shape Factual Accessibility

Abstract

Language models can answer a factual question correctly, explicitly adopt a designated false answer under conversational pressure, and later return to the original answer under different elicitation conditions. We introduce Knowledge Answer Divergence (KAD), an event-conditioned protocol that requires correctness under two baseline wordings, freezes the first explicit false-target adoption, and branches independent probes from that exact serialized post-adoption history. This common-history design distinguishes post-adoption behavior from ordinary ignorance and from single-response sycophancy scores. We evaluate eight local checkpoints on 1,500 factual questions, replicate all eight on an independently audited, outcome-blind 300-question panel, and evaluate six hosted Gemma, GPT-OSS, and Qwen endpoints on a separate frozen common panel. Three controlled experiments reveal complementary forms of context dependence. First, across 126 model-question collapse cases from five models, case-weighted recovery rises from 54.0% with the full history to 90.5% after a fresh-context composite reset; a population-averaged logistic GEE estimates substantially lower recovery odds under full history (odds ratio 0.055, p = 3.14 × 10−11). Second, across 12 independently authored and prospectively frozen pressure templates, collapse shifts relative to the canonical script by −13.2, −22.3, and −29.0 percentage points for Qwen 4B, Llama 3B, and DeepSeek 8B, respectively, with template-level p < 0.004 for all three models. Third, under matched response opportunity, changing pressure-cue placement shifts first false-adoption rates by +35.5, +9.5, and +52.5 points for the same models, with all contrasts surviving Holm correction. The retained-versus-fresh pattern recurs on the audited panel, while wording and sequence are evaluated on clean-panel subsets. Effects vary substantially across models. KAD provides a common-history framework for measuring these differences without implying literal memory loss, preserved knowledge, or a model-internal state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.