acceptodds
Under review as a conference paper at ICLR 2027

Learning from What Goes Wrong: Failure-Aware World Models for Faithful Robot Policy Evaluation

Abstract

Video world models offer a scalable proxy for robot policy evaluation, but successonly training can overestimate policy success when the supplied actions lead to failure. Simply adding failure examples is still insufficient, because the failure space is broad and post-failure object dynamics are difficult to model plausibly, often leading to unrealistic motion or appearance changes that can mislead downstream VLM-based judges. We present a failure-aware framework that combines simulator-verified, scene-matched success–failure trajectories with Semantic-Aware Feedback Forcing (SAFF). SAFF uses feedback and semanticmotion signals to prioritize interaction-relevant prediction errors without changing the flow-matching target or adding auxiliary inference models. We also introduce a six-dimensional benchmark with confidence-aware trajectory evaluation. On simulation and real-world video benchmarks, our model scores 73.17 and 73.31 out of 100, surpassing the strongest baseline by 5.32 and 21.21 points, respectively. Across five policies in each setting, world-model–VLM success-rate estimates achieve Pearson correlations of 0.928 and 0.942 with execution-based rates, while reducing mean absolute error relative to Ctrl-World from 13.0 to 7.0 and from 16.16 to 8.26 percentage points, respectively. These results demonstrate improved prediction fidelity and closer agreement with execution-based policy performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.