acceptodds
Under review as a conference paper at ICLR 2027

CompWorld: Compositional Action Control for Video World Models

Abstract

Interactive video world models increasingly serve as controllable simulators for domains such as gaming, camera-controlled navigation, and robot manipulation, where video generation models must support heterogeneous action types such as keyboard commands, mouse motion, camera poses, and robot-arm trajectories. However, most existing video world models build action-control modules for partic- ular action types, leaving different action types in isolated representations that are difficult to combine reliably without paired multi-action supervision. We propose CompWorld, a framework for compositional action control in video world models. CompWorld keeps the original action streams as precise inputs, while projecting them into a text-aligned unified action space that uses the pretrained text space of a text-conditioned video generation model as a shared interface prior. With action type-specific encoders, learned “not provided” placeholders, and Action-to-Text Fusion, CompWorld enables different action types to condition the video generator through Fused-Text Cross-Attention and compose at inference time. We evalu- ate CompWorld in compositional and separate action-control settings, including two compositional benchmarks—keyboard plus camera pose and robot-arm plus camera pose—that pair heterogeneous controls only at inference, without paired multi-action training annotations. Experiments across keyboard, mouse, camera- pose, and robot-arm controls show that CompWorld improves compositional action control, action following, and video quality over FrameInject, TokenConcat, and GameFactory, and yields a more semantically aligned conditioning interface.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.