OccluCast: Casting Hidden Backgrounds from Multi-View Driving Videos
Abstract
Removing foreground vehicles from driving videos is a fundamental step toward editable, empty-street reconstruction, yet it remains challenging due to severe occlusions, multi-view consistency requirements, and the need to recover geometrically plausible background content. We present OccluCast, a multi-view temporal diffusion framework for vehicle-free driving-scene inpainting. Given synchronized three-view videos and foreground vehicle masks, OccluCast augments a pretrained video diffusion backbone with a lightweight ViewSync branch that injects same-timestamp multi-view attention and temporal attention to propagate background evidence across cameras and time. An empty-street classifier-free guidance strategy further suppresses foreground-vehicle hallucinations by steering the denoising trajectory away from a jointly learned foreground-reconstruction direction. To lift the inpainted videos into an explicit 3D representation, we perform mesh-guided Gaussian distillation, where a LiDAR-derived background mesh filters the supervision loss so that inpainted content is applied only to genuinely occluded regions while real observations supervise visible surfaces. Experiments on the Waymo benchmark demonstrate that OccluCast produces visually realistic, temporally stable, and geometrically consistent empty-street reconstructions, providing a reliable tool for autonomous driving simulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.