The Final Layer Is Rarely Best: Layer-Wise Probing of Visual Foundation Models for 3D Reconstruction and Semantics
Abstract
Visual foundation models (VFMs) are routinely reused as frozen encoders, yet features are often taken from the final layer without testing whether it is appropriate for a dense downstream task. We study layer selection across 14 VFMs and five probes spanning static geometry, texture, joint Gaussian reconstruction, dynamic reconstruction, and semantic segmentation. Our evaluation uses a common Gaussian representation: static novel-view synthesis follows Feat2GS, while Feat4DGS adds a deformation readout for dynamic scenes and dedicated linear probes for semantics. Across the resulting 70 model–task comparisons, a non-final layer strictly outperforms the final layer in 64 cases. The preferred depth is nevertheless task- and model-dependent: the earliest sampled layer is best for texture in 10 of 14 models, whereas semantic and dynamic optima are distributed across the network. No pretraining objective dominates every probe; the best geometry, texture, joint reconstruction, dynamic reconstruction, and semantic scores come from five different model families. Separate deformation readouts also improve mean dynamic-scene PSNR by 4.26 dB over a shared MLP. These results show that layer choice is a first-order experimental variable and provide task-specific guidance for using frozen VFMs in dense 3D and semantic applications.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.