acceptodds
Under review as a conference paper at ICLR 2027

Compositional Diffusion with Hierarchical Latent Cascades

Abstract

Diffusion models are powerful generative tools across images, video, motion planning and beyond. Compositional diffusion extends a pretrained model to horizons longer than those it was trained on by jointly denoising overlapping windows, and a growing body of work keeps the result globally coherent: seamless joints and a consistent layout and style. Yet the composed output often fails the intended global semantics: an object requested once appears in several windows, or in the wrong place. We introduce Hierarchical Latent Cascades (HLC), which routes the global condition down a hierarchy of windows. A forward pass of the pretrained model on a coarse view of the whole output reveals how each part of the condition is distributed over local horizons; we use it to reweight the condition for every child window and pass it down the tree. Each child denoises under its own condition and its parent’s estimate, and overlapping siblings are aggregated level by level into the long-horizon output. HLC is training-free and uses only the model’s own attention, so the same procedure applies to image, video and trajectory models whose condition enters through attention. Experiments on panorama images, long videos and robot motion planning show that it keeps much of the requested global semantics while staying coherent.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.