Do World-Action Models Need Dense World Context? Progressive Sparse Forcing for Action DiTs
Abstract
World-Action Models (WAMs) built on Diffusion Transformers (DiTs) have become a powerful framework for embodied control, generating actions by denoising action tokens conditioned on visual world states. Across diverse visual conditioning layouts, existing WAMs typically ground action prediction through dense action-to-world attention at every DiT layer. This raises an underexplored question: which visual tokens do noisy action tokens attend to during denoising? Analyzing WAMs under mainstream conditioning layouts, we find that dense conditioning often becomes homogenized across depth and task phases, and frequently anchors on task-irrelevant regions rather than action-relevant physical cues. To make visual grounding adaptive, we introduce Progressive Sparse Forcing (PSF), a general action-to-world conditioning framework that converts dense visual-context aggregation into active visual focusing. PSF couples a depth-decaying sparsity schedule with an end-to-end learned action-world router, forcing noisy action tokens to select denoising-relevant visual context from the original training objective. On robotic manipulation, PSF improves closed-loop success on LIBERO and substantially strengthens zero-shot spatial robustness on LIBERO-Plus. On autonomous driving, it improves long-horizon planning performance across 4K-100K training clips from the PhysicalAI-AV dataset. Beyond accuracy, PSF reaches strong performance and reduces inference computation via sparse conditioning. Code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.