Predicting Where Recall Fails: The Fact Matters More Than the Entity in Black-Box Language Models
Abstract
Language models recall some facts reliably and fabricate others, and the standard predictor of where recall fails is entity popularity, which assigns one score to every fact about an entity. We ask whether factual reliability is a property of the entity or of the fact. Holding the entity fixed, we probe multiple numeric facts about the same entity and period in four knowledge domains (Major League Baseball, the National Hockey League, national economic indicators, and corporate finance), covering thirty statistics and seven models from six laboratories. The fact matters more than the entity. In the sports domains, the statistic being asked explains about half of the variance in answer error and the player about 5%. Entity popularity has no held-out predictive value, whereas prose redundancy, the Wikipedia mention rate of a statistic computed before any query, reaches a held-out of 0.41–0.44 and predicts error on statistics withheld from fitting ( of 0.33–0.40), where features of the numbers themselves, such as their dispersion and magnitude, do not. Ordering statistics by prose redundancy places the least mentioned tier lowest in all four domains (fact-level exact permutation test, combined ); deviation correlation falls from +0.894 to +0.163 in baseball and from +0.990 to +0.853 in national statistics. A second pre-query measure, reconstructibility from co-occurring facts, indicates which low-redundancy facts are consistent with inference from headline statistics. For the least mentioned baseball statistics, models fabricate values while refusing fewer than 1% of questions, whereas for a season beyond their training data the non-answer rate rises to 24.8%. For two frontier models, the spread of repeated samples locates this boundary without ground truth, with an AUC of 0.952. The assessment needs only public text and API access to the model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.