acceptodds
Under review as a conference paper at ICLR 2027

Joints, Camera, Action! Co-Generating 3D Motion and Video for Arbitrary Rigs

Abstract

Video generation models can now depict remarkably realistic motion. Yet it is unclear how to use this motion beyond pixels: to animate a specific 3D asset and replay the animation in its 3D scene.For instance, a video model can show a monkey jumping over a crossbar, but it cannot return that jump as replayable 3D motion placed in the 3D scene.Here we present a model that generates, together with the video, 3D motion for any supplied rig and a camera that registers this motion to the scene. Prior joint video–motion models copy billions of parameters from the video model to generate motion. We instead keep 3D motion in a separate model that supports arbitrary skeletons, and connect it to a pretrained video model through a small set of cross-attention layers. This separation lets the motion model first learn how skeletons move from motion-only data, which is easier to obtain and augment than paired video. Additionally, our model learns to locate each joint in the video frames, which lets us recover the camera and place the 3D motion in the scene. To train across many skeletons, we use a generative video editor to place animated 3D assets in realistic scenes, curating the edits so their known motion remains valid. Together with human and robot video, this yields nearly 200K training pairs. Our evaluations demonstrate that our model animates diverse rigs from text more faithfully than video-first pipelines, and its motion responds to the surrounding scene. It also turns video into rig motion, rig motion into video, and enables keyframe-base control of the 3D motion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.