Shot3R: Decoupled Multi-Shot Reasoning for Streaming 4D Human-Scene Reconstruction
Abstract
Multi-shot video splices shots taken at different viewpoints over a continuous timeline. A shot change jumps the viewpoint, breaking identities and spatial reference at the shot change. Reconstructing people and scene from such video requires spatial and identity consistency across the changes. Most existing methods resolve cross-shot ambiguities by optimizing all shots and the scene jointly, once the complete sequence has been reconstructed. Unlike such offline methods, streaming reconstruction can only infer such relations from current and past observations, with no global information and only a history from the previous viewpoint. To address this problem, we present Shot3R, a streaming framework for 4D reconstruction of multiple people and scene from monocular video. At each shot change, Shot3R infers the transform between consecutive shots in latent space, from the boundary frame's features and the accumulated history, placing the new shot in the shared world frame. Identities are associated across the shot change. The transform is applied to the whole new shot, freeing it from the influence of the previous shot. On extensive multi-shot data, Shot3R delivers marked improvements in human reconstruction, camera pose estimation, and cross-shot identity consistency, largest under large viewpoint changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.