Knowing What You No Longer Know: Calibrated Belief Validity for Persistent Perceptual World Memory
Abstract
Persistent world memories let embodied agents answer questions about objects they saw minutes or hours ago. Every such system that reports a confidence reports descriptive uncertainty, how sure it is what an object is, and none reports validity uncertainty, how sure it is that a belief is still true, as a fitted or evaluated quantity. The consequence is systematic and invisible to current benchmarks: asked where an object is after it moved off-camera, models from 14B to 235B answer with the stale location on every such real-kitchen item and never once flag it as stale, at confidently-wrong rates of 1.000 to 0.970 and still 0.879 when told that declining is acceptable, while correctly declining about objects they never saw. We prove that any confidence which is a function of object class and elapsed time is at exactly chance on an elapsed-time-matched pair, and we introduce STILL-Bench, a benchmark with derived epistemic ground truth over real changing kitchens and controlled ones, built from such pairs. Fitting an interval-censored hazard to 9,120 real observation gaps, we find that validity is predictable, at AUROC 0.765 against 0.587 for fitted recency, and that the signal is almost entirely where the object was last seen: expiry runs from 0.029 on a counter to 0.614 in a drawer. A place-conditioned validity model separates stale from fresh beliefs where every published time-based heuristic is provably at chance, at PVD 0.710 (p = 0.029), answers 98.0% of queries at 5% risk or less where recency answers 61.3%, and ranks places well enough to be the best budget-limited re-observation policy we tested on procedurally generated houses. Three cautions follow: descriptive uncertainty does not predict validity and is anti-predictive on pooled pairs; a simulator whose injected changes lack the real place regularity inverts a validity evaluation; and the fitted hazard does not transfer between two real datasets, because a belief formed by seeing an object in an enclosure and one formed by putting it there carry opposite hazards.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.