acceptodds
Under review as a conference paper at ICLR 2027

Infusing Motion Intelligence into Structured 3D Worlds for 4D Scene Generation

Abstract

Generating a 4D scene from a single image requires endowing a structured 3D world with meaningful and coherent dynamics. Existing image-to-4D methods predominantly introduce dynamics through video generation and subsequently reconstruct 3D scenes. While video models provide strong motion priors, their dynamics are formed in 2D space, making scene geometry, physical constraints, and fine-grained motion control difficult to explicitly access and verify. In contrast, explicit 3D representations provide a structured world in which geometry, physical conditions, and motion states can be directly represented and manipulated, but lack the rich motion intelligence required to translate semantic instructions into meaningful dynamics. We propose SceneWeaver, a 3D-native 4D generation framework that bridges this gap through a Coordinator-Executor-Verifier architecture. Given an image and a prompt, a coordinator interprets the motion intent and translates it into structured conditions describing where the subject should move, how it should move, and how the environment should affect its motion. Two specialized executors then generate the corresponding global world-space trajectory and local articulated motion within the explicit 3D scene. Specifically, a potential-augmented Neural ODE, which is geometry- and physics-aware, generates the global trajectory by jointly modeling interpreted motion control, scene geometry, and environmental effects, while an LLM-based motion planner produces fine-grained articulated motion from structured motion descriptions. Crucially, the explicit 3D representation also enables numerical verification of intermediate motion states. A verifier identifies geometric, physical, and motion-related failures, localizes root causes and enables targeted repair through coordinator. This establishes a closed-loop generation in which 3D structure serves not only as the intermediate results, but also as the basis for verification and recovery. Experiments demonstrate that SceneWeaver shows promising controllability, motion quality in generated 4D scenes, offering an alternative paradigm for image-to-4D task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.