Decision-Relevant Representations under Rare Observations
Abstract
World models and predictive representations decide what to remember by how well the past predicts the future. This objective has a blind spot: a cue that is rare, unpredictable from the past, and yet decides the action—a traffic light that turns red—contributes almost nothing to an average predictive loss and is the first distinction compression discards. We show that the right currency for such cues is evidence, not frequency. We introduce coverage-calibrated Bellman quotients (CBQ), a latent-state learner that merges two histories only when their counts certify that no action consequence separates them, and returns the most compact code that passes this certificate. CBQ comes with a complete finite-sample account. Its supported -error splits into approximation bias and statistical error ; it keeps every action-flipping cue whenever , recovers the exact Bellman quotient—discarding all decision-null nuisance—under a separation condition, and turns both into fresh-rollout regret bounds. A matching obstruction shows that no learner can identify a rare decisive cue when , so in the one-stage case the evidence requirement of CBQ is tight up to logarithms. Across 288 random POMDPs, predictive objectives keep the rare cue in at most of instances at any sample size, while CBQ keeps it in all of them with a code exactly as compact as the Bellman quotient; all 4,320 runs switch on within one decade of the theory's coverage statistic , and the certificate is never violated. From pixels, next-frame and self-predictive world models lose a rare traffic light in 107 of 108 runs, whereas CBQ on a learned patch tokenizer keeps exactly the light's patch and discards the watermark entirely.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.