Candidate-Dependent Margins in Latent World-Model Evaluation
Abstract
Candidate actions change the measured margins between fixed latent world models, even when their top-ranked predictor remains unchanged. We quantify this dependence through paired candidate sets at fixed checkpoints, initial states, goals, horizons, and candidate counts. Among seven predictors sharing an encoder on 64 Push-T families, all six mean gain advantages over a post-hoc logged-data comparator contract by –. Their raw prediction-error advantages also contract by –, so this sensitivity extends beyond reference normalization. These six paired gaps share one comparator and do not constitute independent replications. Twelve of 21 post-hoc score-response contrasts survive a studentized sign-flip test with Holm correction; paired tests retain the same findings. Separately, a public DINO-WM reproduction gains in persistence-relative score despite an increase in prediction error. Family-level accounting separates the opposing model-error and reference-error terms. A prospective PointMaze study identifies worse available task opportunity without resolving a change in within-set regret. We provide a reproducible audit of prediction error, reference difficulty, available opportunity, selected cost, regret, and cost discrimination. These measurements quantify candidate-dependent comparison margins and delimit their implications for action selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.