Separating Action from Prediction in World Action Models
Abstract
World action models (WAMs) learn latent variables from observations and use them both to predict how the scene evolves and to generate actions. Prediction has to account for the whole scene, whereas the actions depend on only part of it. However, WAMs fitted to demonstrations provide no guarantee of generating actions only from the latent variables on which the demonstrated actions depend, termed action-relevant latent variables, because a model may imitate equally well while entangling them with nuisance variables. We therefore separate what the actions may depend on from what prediction needs, and impose a sparsity constraint only on the estimated structure between latent states and actions. We prove that, under this constraint, the latent variables used to generate actions are an invertible function of the action-relevant ones alone, and the true structure is identified up to a permutation of the latent variables. These guarantees require neither sufficient temporal dependence nor assumptions on how unobserved regimes evolve, and hold even when the action-relevant latent variables are statistically dependent on the rest. We approximate the constraint in a pretrained WAM by adding a sparse routing path to its action head, while the world-modeling objective stays unchanged. Synthetic experiments support the theory, and on simulated and real robots, the routing path improves aggregate closed-loop success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.