acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Candidate Replacement under Imperfect World Models via Baseline-Relative Learning

Abstract

World-model errors can make candidate scores unreliable. A higher predicted return does not always justify replacing the agent's incumbent action. We propose Baseline-Relative Value Learning (BRVL), a post-training selector for frozen world-model agents. BRVL learns candidate–incumbent return differences from paired offline simulator rollouts under a common baseline continuation. Its objective combines relative-return regression with a candidate-set loss for forgone gains and harmful replacements. At deployment, BRVL replaces the incumbent only when an alternative has a positive calibrated lower score. We evaluate BRVL with DreamerV3 and TD-MPC2 on DMControl, Meta-World, and ManiSkill2. BRVL improves mean closed-loop performance and frozen-set net gain over matched-supervision regression baselines in all six settings. The gain–harm trade-off varies across settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.