acceptodds
Under review as a conference paper at ICLR 2027

Stable Self-Evolving World Action Models via Self-Perplexity Aware Learning

Abstract

While self-evolution shows great promise for improving model performance without human demonstrations, how to self-evolving World Action Models (WAMs) that jointly model actions and videos remains unexplored. However, directly fine-tuning WAMs on self-generated trajectories induces a pronounced V-shaped curve where performance drops sharply during early training and then gradually recovers. This may stems from the distribution shift between self-generated trajectories and pre-training human demonstrations. To address this challenge, we propose SPASE, a multi-round self-evolution framework for WAMs. SPASE first introduces an Execution-Grounded Dual Verification module to filter out low-quality data with incorrect actions or inconsistent videos. To alleviate the adverse effects of distribution shift, SPASE introduces Self-Perplexity to measure data divergence from the model’s parametric knowledge. Using Self-Perplexity, SPASE adaptively downweights highly divergent data to stabilize training and imposes a deterioration penalty on previously learned data to mitigate catastrophic forgetting. Extensive evaluations across robotic manipulation benchmarks demonstrate the superiority of SPASE. Our code is available at https://anonymous.4open.science/r/stable_wam_evolution/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.