Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting
Abstract
Video re-shooting re-renders a monocular video of a dynamic scene along a user-specified camera trajectory. A dominant paradigm explicitly provides the target geometry by lifting the source video into a 4D point cloud using per-frame depth, and rasterizing this representation along the target trajectory to obtain point-cloud renders as geometric conditions for trajectory-controlled generation. Because the point cloud render and the source video are both handed to the network as visual conditions, their potentially conflicting cues persist throughout denoising — a tension we refer to as a trust dilemma — which may affect trajectory control, the 3D consistency of dynamic content, or visual quality, particularly on data outside the training distribution. We argue that because a point-cloud render is pixel-aligned with the target view, it does not need to be continuously provided as a conditioning signal throughout denoising. We propose , which injects the render directly into the initial noise of flow matching, so that generation no longer departs from standard Gaussian noise but from a geometry-bearing noise distribution that we term the point cloud rendered manifold, leaving the source video as the only visual condition. The render is thus used exactly once, and the network is never asked to learn how to read it; in subsequent denoising steps the model can focus on the source video. On our benchmarks, Manifold4D attains the best camera-control accuracy on every metric, lowering rotation error by 25%–27% and translation error by up to 32% over the strongest baseline, while matching it in video fidelity and attaining the best overall photometric quality and 3D consistency of dynamic content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.