acceptodds
Under review as a conference paper at ICLR 2027

Do World-Action Models Need Dense World Context? Progressive Sparse Forcing for Action DiTs

Abstract

World-Action Models (WAMs) built on Diffusion Transformers (DiTs) have become a powerful framework for embodied control, generating actions by denoising action tokens conditioned on visual world states. Across diverse visual conditioning layouts, existing WAMs typically ground action prediction through dense action-to-world attention at every DiT layer. This raises an underexplored question: which visual tokens do noisy action tokens attend to during denoising? Analyzing WAMs under mainstream conditioning layouts, we find that dense conditioning often becomes homogenized across depth and task phases, and frequently anchors on task-irrelevant regions rather than action-relevant physical cues. To make visual grounding adaptive, we introduce Progressive Sparse Forcing (PSF), a general action-to-world conditioning framework that converts dense visual-context aggregation into active visual focusing. PSF couples a depth-decaying sparsity schedule with an end-to-end learned action-world router, forcing noisy action tokens to select denoising-relevant visual context from the original training objective. On robotic manipulation, PSF improves closed-loop success on LIBERO and substantially strengthens zero-shot spatial robustness on LIBERO-Plus. On autonomous driving, it improves long-horizon planning performance across 4K-100K training clips from the PhysicalAI-AV dataset. Beyond accuracy, PSF reaches strong performance and reduces inference computation via sparse conditioning. Code will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.