acceptodds
Under review as a conference paper at ICLR 2027

World-Grounded Camera Trajectory Generation

Abstract

Driven by advancements in 3D world reconstruction, the demand for world-aware video generation and navigation is rapidly growing, where generating controllable camera trajectories is the prerequisite for ensuring 3D consistency. While existing methods show promise, they heavily rely on character-driven motions or single RGB-D images, leaving text-driven trajectory generation in unconstrained 3D point clouds largely unexplored. In this paper, we present WorldDirector, a novel framework that generates world-grounded camera trajectories directly from 3D point clouds and text prompts. To achieve precise spatial grounding, we propose camera-aware world tokens that explicitly encode the observer’s initial perspective by concatenating point-wise geometry features with 3D coordinates transformed into the starting camera’s local coordinate system. This representation allows a Diffusion Transformer (DiT) to effectively map semantic intent onto geometric constraints through cross-attention layers, ensuring that the generated trajectories are both physically plausible and strictly consistent with the 3D scene. To support this task, we introduce WorldTraj, an extensive dataset consisting of roughly 136K triplets of 3D worlds, camera trajectories, and text prompts spanning 90K diverse environments. Extensive experiments demonstrate that WorldDirector successfully generates camera trajectories that are highly consistent with both the 3D world geometry and the semantic instructions in the text prompts. Code and models will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.