acceptodds
Under review as a conference paper at ICLR 2027

Better Samples, Wrong Odds: When Prediction Rewards Distort World-Model Probabilities

Abstract

Prediction rewards post-train world models to produce better samples; we study when they preserve the probabilities a planner needs across futures and when they distort them. A prediction reward, which we call a point reward, compares one generated future with one observed outcome. Under such a reward a model can score below the truth's own risk only with wrong probabilities, which gives a checkable diagnostic: for any metric, the excess gain—the part of the gain that lies below the truth's own risk—lower-bounds the Wasserstein-1 distance to the truth. In PushT, where all futures can be enumerated, about 72% of the gain from point-reward training is excess gain. On RLVR-World's RT-1 video world model, point rewards lower pixel error but give a worse held-out joint energy score than an energy-score reward from the same base, data, and budget in all five seeds, by 2.1% of Base's score at 200 updates and 7.2% at an exploratory 800. The distortion reaches decisions when action rankings depend on how probability is split: in an image-conditioned MiniGrid task, it costs 10.3 points of success, and continuing with a proper score recovers most of it. Proper scoring alone is not sufficient: the update rule also matters. Under idealized GRPO standardization, energy-score and point-reward updates coincide for binary outcomes, while unbiased leave-one-out credits keep the truth a fixed point. We therefore recommend reporting a proper score beside sample metrics and training proper rewards with unbiased credits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.