PoseSpecGS: Pose-Aligned Multimodal Gaussian Splatting with Spectrum-Response Rendering
Abstract
Multimodal novel-view synthesis aims to reconstruct a unified scene representation from images acquired by sensors with distinct image-formation characteristics. However, modality-dependent appearance and inaccurate camera poses hinder learning a compact, reusable representation. We present PoseSpecGS, a pose-aligned Gaussian splatting method with spectrum-response rendering for multimodal novel-view synthesis. Inspired by sensor spectral responses, we first introduce Spectrum-Response Rendering, which synthesizes modality-specific images from alpha-composited shared Gaussian spectral features using compact modality-channel response operators. To address pose inaccuracies, we further introduce Cross-Modal Alignment, which registers auxiliary-modality cameras to an RGB structure-from-motion backbone and jointly refines their poses and scene parameters using RGB-guided reprojection consistency, without known rig extrinsics or synchronized frame correspondences. The representation supports static and time-conditioned multimodal reconstruction. Moreover, shared response operators pretrained across scenes enable zero-shot one-to-multimodal transfer under RGB-only supervision. Experiments demonstrate state-of-the-art reconstruction quality on the evaluated benchmarks, with reduced model storage and fewer training iterations. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.