Imaging the Radar: Recovering 4D Radar Tensors from Images with a Latent Diffusion Bridge
Abstract
Millimeter-wave radar's all-weather robustness and Doppler velocity sensing make it increasingly important for autonomous perception. However, the scarcity of high-quality 4D radar data limits the development of radar perception systems. Images provide a scalable source for radar data expansion: camera sequences are abundant and inexpensive to acquire, preserve scene structure shared with radar, and support scene-level editing. We therefore introduce image-to-4D radar tensor generation and propose the Radar-aligned Visual Prior Bridge (RaViP-Bridge) to recover complete radar tensors from images. However, direct translation faces a substantial modality gap. Images record scene appearance on a 2D perspective plane, whereas native 4D radar tensors encode spatial and Doppler responses amid noise and clutter. To narrow this gap, we use visual foundation models to extract geometry, semantic, and motion cues from image sequences and organize them into a radar-aligned visual-prior tensor. Given the high dimensionality and noise of native radar tensors, we use dual VAEs to compress the visual-prior and radar tensors into compact latent spaces. We then formulate latent translation as an endpoint-constrained diffusion bridge from the visual-prior latent to the radar latent. On K-Radar, we find that RaViP-Bridge achieves high-fidelity 4D radar generation and improves downstream detection through data augmentation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.