HashMod: Plug-and-Play Patch-Wise Hash-Guided Refinement for Feed-Forward 3DGS
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) has emerged as an important approach to efficient and generalizable 3D scene reconstruction, typically decoding image features into Gaussian primitives in a single forward pass. However, reconstruction quality may degrade on unseen scenes with different geometric structures, texture patterns. Fine-tuning only the Gaussian head offers a low-cost adaptation strategy, but its ability to correct local geometric and appearance errors can be limited by insufficient emphasis on relevant backbone features.To address this limitation, we propose HashMod, a lightweight modulation module that uses patch-wise binary query–key matching to adaptively select both the number and identity of backbone feature levels contributing to each local correction. HashMod aggregates the selected feature-derived offsets to multiplicatively refine the initial Gaussian predictions. Across three public benchmarks and cross-dataset settings, HashMod achieves competitive or improved reconstruction quality across AnySplat, YoNoSplat, VGGT-, and Depth Anything 3, while introducing only approximately 8M additional parameters and reducing training time by approximately 41–42%. Furthermore, we conduct fine-tuning experiments on two self-collected crop datasets, each comprising 1,000 videos. The adapted feed-forward 3DGS models improve reconstruction quality on real-world crop scenes, demonstrating the practical applicability of HashMod beyond standard benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.