Retrieval Apophenia: How Growing Candidate Pools Make New Events Look Familiar
Abstract
Streaming event induction asks whether an arriving document belongs to an event already in memory or starts a new one. Most systems answer this question by retrieving historical candidates, taking the candidate with the highest compatibility score, and comparing that score with a fixed threshold. We show that this maximum rule fails in a specific and predictable way: as memory grows, the maximum is taken over more candidates, making a genuinely new event increasingly appear familiar. We call this phenomenon candidate-pool false familiarity. Prior work has studied threshold drift in streams and score inflation in large candidate galleries, but these effects have not been connected within streaming event memory, where an erroneous merge changes the state against which subsequent documents are judged. We measure this effect on annotated non-coreferent candidates and distinguish between two possible explanations. If the drift were merely an uncalibrated score scale, normalising each query by its candidate pool should eliminate it. However, this removes only a small portion of the effect, indicating that the decision must examine the candidates themselves rather than rely on a single score. We therefore propose C2E-Judge, which keeps the retriever, candidate budget, and memory fixed while allowing a language-model judge to inspect the entire bounded pool. It separates the two error types: adding candidates improves repeated-event coverage with nearly no increase in false merges, whereas replacing score thresholding with whole-pool judging removes most false merges. This visibility effect persists under a stricter prompt, a second judge endpoint, and a second dataset.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.