acceptodds
Under review as a conference paper at ICLR 2027

Layout2Video: Persistent 3D Layout Control for Indoor Video Generation

Abstract

Text- and image-conditioned video diffusion models produce visually compelling videos, but their inputs do not explicitly specify scene-wide 3D object arrangements. Maintaining these arrangements as the camera reveals new regions therefore remains challenging. We present Layout2Video, a framework for generating indoor videos from 3D semantic layouts and text prompts without task-specific training. The framework first renders view-specific semantic and geometric conditions from the layout. A pretrained multi-view diffusion model uses these conditions and an initial reference image to generate candidate views, from which sparse anchors are selected. A video diffusion model then synthesizes segments between consecutive anchors, and the segments are concatenated into a single video. All generative components remain frozen. We also construct a benchmark containing 555 room layouts and introduce an instance-level evaluation protocol for object presence, layout adherence, and appearance consistency. Extensive experiments demonstrate that our method effectively preserves the specified object arrangement and appearance throughout generated videos, particularly as previously unseen regions become visible.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.