MoE4D: Mixture of Spatio-Temporal Experts with Frequency-Modulated Router for LiDAR Novel View Synthesis
Abstract
Although recent progress in 4D scene representations significantly advances LiDAR novel view synthesis, it still remains inconclusive which spatio-temporal representation performs better across diverse regions and time-variant motion patterns. NeRF-based methods encode continuous 4D implicit field with superior parameter efficiency, while struggling in long-term dynamics and occlusion scenarios. In contrast, Gaussian Splatting-based methods enable fast real-time rendering but commonly require complicated ray-object intersection modeling, dense trajectory labels of dynamics, and explicit storage of Gaussian primitives. In this paper, we design a novel differential LiDAR simulation method based on Mixture-of-Experts, called MoE4D, to effectively combine various spatio-temporal cues within a unified rendering pipeline. Specifically, multi-planar, hash grid, and object-centric GS serve as main expert candidates, providing multi-granularity geometry and deformation priors. We also observe that previous works mostly overlook spectral discrepancies within LiDAR sequences. To mitigate this issue, a frequency-modulated context router is designed based on the Fourier analysis of input 4D signals. Furthermore, a dense-to-sparse expert distillation strategy is developed to maintain high-fidelity reconstruction quality with significantly reduced computational budgets. Extensive experiments on KITTI-360 and nuScenes demonstrate state-of-the-art performance of our MoE4D. Codes will be released upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.