A Downstream Metric Can Be Non-Diagnostic for Predictive-Fit World-Model Selection
Abstract
World models are often selected by held-out predictive fit and then judged by downstream decision quality. Yet a downstream metric can itself be non-diagnostic of the selector it evaluates: a weak or null comparison cannot distinguish a poor selector from an insensitive evaluator if the metric does not respond to known increases in decision-relevant information. In a controlled finance testbed with fixed model families and policy pools, the inherited realized-path regret endpoint fails this test: a true-dynamics expected-return selector does not separate from the pre-specified within-family random-model calibration. We repair the estimand to expected net return and replicate its information sensitivity on a fresh population block under family-wise controls before examining the selector comparison. Under the gate-passing endpoint, 0/8 pre-specified finance slices show a multiplicity-controlled detectable advantage of held-out one-step fit over uniform model identity after Holm correction; this is non-detection, not equivalence. A post-hoc order-statistic diagnostic finds no slice with positive excess hindsight-best-family headroom relative to references matched to the observed number of distinct decision classes. This exploratory result weakens the explanation that one substantially better candidate was merely hidden among poor models. The contribution is a validity-gated selector-inference procedure that treats information sensitivity as a prerequisite, demonstrated through a preserved failure–repair–replication sequence rather than a new world model or selector.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.