acceptodds
Under review as a conference paper at ICLR 2027

When Retrieval Fails: Structured Partial Information in Language Models

Abstract

Language-model evaluation typically treats factual retrieval as binary: correct or not. Human memory research instead distinguishes successful recall from tip-of-the-tongue (TOT) states, where partial information about an unretrievable target remains accessible. We ask whether language models show a computational analogue, and state our result up front: we do not establish that models retain target-specific information after retrieval failure. We establish a weaker phenomenon – candidate partial-retrieval failures show greater aggregate accessibility to answer attributes than matched low-knowledge failures, via a Partial Retrieval Score (PRS) over independent attribute probes. Across three open-weight 7–14B models and two natural-item samples (up to 857 items), this aggregate effect replicates robustly () and survives restriction to label-stable items. But every attribute-level lead we identified – geographic region, then character length – fails independent replication, a forced-choice specificity probe, and pre-cue hidden-state probing, and run-to-run label agreement is only under quantized greedy decoding. A causal cue-intervention experiment shows a correct-value cue helps retrieval equally in both groups (the -vs- interaction is null) – evidence that cue content is useful once supplied, not that it was already present internally. We treat the aggregate PRS effect as the paper's one claim to survive scrutiny, and attribute-level or internal-state claims – including our own – as requiring more than a single pilot to trust.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.