Silence Is Not Absence: A Probe Null Requires More Than Format Competence
Abstract
A probe that asks what a language model memorised reports a null when the model answers at chance, and that null is evidence of absence only if the probe could have shown the fact. Two separate conditions have to hold, and checking the first does not establish the second. The first is that the model can use the readout. Asked thirty questions with no domain content, of the form “which one is a fruit, an apple or a hammer”, each in both option orders so that a fixed letter preference scores exactly one half, seven of nineteen open-weight models cannot do it. The failures are not the small models: every rung of the Pythia ladder fails from 0.16B to 6.9B, and Llama-2-7B fails while Llama-3.2-1B passes at a sixth of its size; inside one laboratory OLMo-1-7B scores 0.517 and OLMo-3-7B scores 1.000. Two of the seven answer a yes/no frame at 0.983, consistent with a label-binding failure rather than an inability to answer, and the excluded models cannot be read even when the answer is printed in the prompt, at 0.507 and 0.535. The second is that information of the kind being probed can reach that readout, which no control built from questions a model already knows can establish. We test it by writing invented dated facts into training text at four exposure counts, in three models from three laboratories. Under the form in which they were learned every model separates the true value from the probe's own foil, increasingly with exposure and against flat untrained controls, while under five forced choices — relabelled, list-free, and with demonstrations — the same facts remain at chance in every model. Interface validity does not imply construct validity. Turning the instrument on the question that motivated it, whether small open-weight models memorise market history closely enough to contaminate a backtest, the probe that scores about 0.90 on which company a ticker belongs to scores about 0.50 on every date-specific fact about that company, and we exclude effects larger than 1.75 points in the direction memorisation predicts. By the second condition, that bounds what a forced choice retrieves, not what the models hold. We release the panel, the probes, the estimators and the power calculator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.