VideoPort: Streaming Photorealistic Human Telepresence from a Single View via Real-Time Video Diffusion
Abstract
Human telepresence aims to enable remote viewers experience a live performance from freely chosen viewpoints, however, achieving this from a single RGB camera requires inferring surfaces and appearance that the camera cannot observe. Exist- ing approaches face a fidelity–realism tradeoff. Feedforward 3D reconstruction is geometrically consistent, but often produces coarse appearance. Video diffu- sion synthesizes realistic detail, but when conditioned on its own history, drifts from the person being captured. We present VideoPort, a two-stage framework fixes geometry via feedforward 3D reconstruction, then injects realistic detail us- ing causal video diffusion refinement. Reconstructed geometry guides novel-view synthesis, while surface-aligned observations provide additional appearance con- ditioning during training. To address error accumulation through generated his- tory, we propose deep self-rollout supervision (DSR) with paired performance targets to train the refiner to recover from its own imperfect predictions. As a result, VideoPort preserves performance-specific appearance throughout minutes- long sequences and achieves state-of-the-art reconstruction quality on 3 bench- marks. We further deploy it as a live system that performs continuous, real-time novel-view synthesis from a single consumer RGB camera.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.