acceptodds
Under review as a conference paper at ICLR 2027

CamCanvas: Absolute Camera Placement in Video World Models

Abstract

Interactive video world models generate long sequences under camera commands and sustain extended exploration of a generated world. Their camera interface comes from relative pose encodings: a command specifies motion with respect to the current context window. As a result, when the camera is commanded back to a place visited long ago, these models fail to return to the commanded viewpoint. We present CamCanvas, which stores the model’s own generated content in a panoramic canvas, retaining the first observation of each location. When the camera returns, the model is conditioned on the rendering of this canvas at the commanded pose. The model thus receives the target view directly, so it has neither to infer its absolute pose from relative motion nor to recover that view from a limited context window. The return test assesses absolute camera placement by commanding the camera away and back and estimating, from the frames, the viewing-angle difference between the return and the model’s own first visit, with no external alignment. Return success rises from 0% to 91% on 90◦ returns and from 18% to 95% on real video. CamCanvas also generalizes to a video world model with a different architecture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.