acceptodds
Under review as a conference paper at ICLR 2027

AnyThermalSplat: Scaling Pose-Free Feed-Forward Thermal Scene Reconstruction

Abstract

Thermal 3D reconstruction typically relies on known cameras and per-scene optimization, while the scarcity of multi-view thermal data hinders scalable feed-forward reconstruction. We present AnyThermalSplat, a pose-free feed-forward framework that predicts renderable thermal Gaussian scenes without external camera parameters or test-time scene optimization. It adapts RGB-pretrained geometric priors through low-rank updates and residual feature bridges, with a frozen RGB teacher providing geometric supervision alongside thermal reconstruction and held-out novel-view objectives; inference requires thermal images only. To expand supervision, we introduce ThermalVerse, a multi-domain RGB–thermal dataset, and a geometry-grounded thermalization pipeline that fits shared thermal appearance on fixed RGB-reconstructed geometry. Together with curated public data, these resources form a 500-scene training pool. On in-domain scenes, AnyThermalSplat achieves 26.76 and 28.48 dB PSNR with 6 and 12 context views, respectively, outperforming our matched thermally adapted AnySplat-IR baseline by 2.84 and 1.77 dB. This baseline is trained on the same thermal data under the same adaptation protocol, providing a controlled comparison beyond directly applying the original RGB model. We also conduct a controlled scaling study on nested 100–500-scene subsets under a fixed update budget and find substantial gains in both in-domain and out-of-domain reconstruction as the training set grows from 100 to 300 scenes. Beyond this regime, out-of-domain performance largely saturates, while in-domain reconstruction retains an overall benefit from larger training sets. We will release the data and code, along with model checkpoints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.