acceptodds
Under review as a conference paper at ICLR 2027

DreamChaser: A Self-Improving Framework for World Action Models via Video-guided Action Refinement

Abstract

World Action Models (WAMs) built on pre-trained video generators learn video prediction and action generation under asymmetric supervision. The video branch inherits dense spatiotemporal supervision from large-scale video pre-training, whereas the action branch relies on comparatively limited paired action annotations. We hypothesize that this asymmetry in data scale contributes to a capability gap: a WAM can predict plausible task outcomes that its generated actions fail to realize. We introduce DreamChaser, a self-improving framework that uses the model’s video predictions to guide action refinement and acquire additional action supervision through interaction. A lightweight refiner, trained solely on the original demonstrations, uses predicted visual outcomes as targets for correcting local execution deviations. DreamChaser selectively aggregates the resulting rollout data and combines them with the original demonstrations to fine-tune the WAM. This process converts video-guided corrections into additional action supervision, forming an iterative loop of interaction and policy updates without further human demonstrations or an external expert during rollout. On LIBERO, DreamChaser fine-tuning increases FastWAM-joint’s average success rate from 97.3% to 97.9% when evaluated without the refiner. On LIBERO-plus, three rounds of self-improvement increase the mean task success from 72.1% to 80.1% across ten tasks. On a real AgileX PiperX, DreamChaser improves the FastWAM-joint average success score on cup sleeving, plate placement, and bottle straightening from 27.0% to 41.7% after two online rounds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.