Calibrated Candidate Replacement under Imperfect World Models via Baseline-Relative Learning
Abstract
World-model errors can make candidate scores unreliable. A higher predicted return does not always justify replacing the agent's incumbent action. We propose Baseline-Relative Value Learning (BRVL), a post-training selector for frozen world-model agents. BRVL learns candidate–incumbent return differences from paired offline simulator rollouts under a common baseline continuation. Its objective combines relative-return regression with a candidate-set loss for forgone gains and harmful replacements. At deployment, BRVL replaces the incumbent only when an alternative has a positive calibrated lower score. We evaluate BRVL with DreamerV3 and TD-MPC2 on DMControl, Meta-World, and ManiSkill2. BRVL improves mean closed-loop performance and frozen-set net gain over matched-supervision regression baselines in all six settings. The gain–harm trade-off varies across settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.