Calibrating Imagination by Decision Regret
Abstract
Action-conditioned world models can replace some camera updates, but visual prediction error alone does not measure the consequence of acting on an imagined observation. DREAM-CAL estimates policy-conditioned decision regret from paired continuations and selects a refresh horizon using a block-calibrated upper bound. Across 37 tasks at nominal 25% visual refresh, DREAM-CAL achieves 1,575/2,220 (70.95%) versus 1,392/2,220 (62.70%) for a visual-uncertainty head, a difference of 8.24 percentage points from the integer counts. In a separate segment audit, harmful-event AUROC is 0.790 versus 0.665, while the visual head retains lower expected calibration error. Gains vary by task family, and four of 37 task counts favor the visual head. We evaluate closed-loop success, event ranking, regret-bound coverage, distribution shift, and inference cost. Under dynamics shift, coverage with the original calibration offset falls from 87.92% to 71.25%, and mean decision computation takes 452 ms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.