MIXTUREPRIOR: ACTION-ROUTED TRANSITION PRIORS FOR LATENT WORLD MODELS
Abstract
Latent world models train policies on imagined rollouts, in which every future state is sampled from a learned stochastic transition prior. Although this prior is conditioned on the action, a single set of parameters must fit the next-state distribution for every action. A correction that helps one action can therefore hurt another, and such opposing corrections cancel in the shared parameters, so the prior can stall while individual actions remain under-fitted. Our key idea is that the prior should offer several candidate corrections and that the action should decide where each one is used. We propose MIXTUREPRIOR, which predicts multiple categorical transition components from the backbone's action-updated feature and combines them with an action-only gate. Because it replaces only the stochastic prior, MIXTUREPRIOR plugs into recurrent and Transformer-based world models without further changes. We analyze this design through the correction value of a component and show that, when a component helps some actions but hurts others, action-dependent weighting achieves a strictly larger first-order improvement than any fixed mixture. We integrate MIXTUREPRIOR into DreamerV3 and TWISTER and evaluate it on 26 Atari-100k games and 20 DeepMind Control tasks. With everything else fixed, four components improve the aggregate mean over a single-component prior in all four backbone–benchmark combinations. In a trained model, we further find components whose correction value changes sign across actions, which is the condition under which our analysis predicts action routing to help.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.