acceptodds
Under review as a conference paper at ICLR 2027

Video Diffusion Rendering for 3D Asset Insertions in Driving Simulation

Abstract

The growing availability of reusable 3D assets and reconstructed real-world scenes is opening new opportunities for scalable and controllable driving simulation. A key capability is asset insertion, where external objects are placed into existing scenes to create diverse and interactive scenarios. However, geometrically rendered insertions often remain visually implausible, suffering from appearance mismatch and limited scene interaction, especially when assets originate from different scenes. While frame-wise harmonization can improve visual realism, it often introduces temporal inconsistency in video outputs. We address this naive-to-real rendering problem with a one-step video diffusion harmonizer that maps naive insertion renderings to photorealistic and temporally consistent driving videos. The harmonizer is obtained by fine-tuning a pretrained video diffusion model, with temporally agnostic latent injection introduced to recover spatial details lost during 3D VAE decoding. To effectively train the harmonizer, we construct DriveInsert-3K, an offline paired video dataset built with a two-stage alignment procedure to provide clean naive-insertion/real-video supervision. At inference time, our model takes a naive insertion video from an upstream insertion pipeline and produces a harmonized video in a single denoising step. Experiments show that our approach improves visual realism and temporal consistency across diverse driving scenes and inserted assets while enabling efficient single-step inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.