acceptodds
Under review as a conference paper at ICLR 2027

SceneQuery: Target-Queried Scene Memory for Visuospatial Video Generation

Abstract

Visuospatial video generation requires synthesizing videos that satisfy target-object or camera-trajectory instructions while remaining consistent with an observed environment. Although existing methods condition generation on scene observations, how to retrieve task-relevant scene evidence for individual output positions remains less explored. We introduce SceneQuery, a task-conditioned scene retrieval framework in which task instructions guide not only what to generate, but also what scene evidence to retrieve. SceneQuery constructs global and spatial memory from RGB observations and derives a query for each output token using target-object semantics and token coordinates for grounding, or the corresponding target-camera ray for navigation. Unlike queries derived from evolving target features, these task-derived queries remain fixed throughout denoising, providing a stable, output-aligned retrieval signal. The retrieved evidence conditions video generation, while the memory and retrieval modules are jointly optimized through the generation objective without additional memory supervision or explicit 3D reconstruction. Controlled query-source comparisons support task-derived retrieval across both tasks. On ScanNet++, selected configurations improve target-view reconstruction for navigation and reduce geometric error for grounding relative to a matched context-guided baseline, while the primary joint grounding-success rate remains unchanged.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.