Visualizing Dense Neural Representations
Abstract
Modern dense-feature extractors such as DINOv3 or Perception Encoder achieve strong performance on a variety of spatially fine-grained tasks. For a qualitative understanding, the features are usually mapped onto the three first principal components which are treated as RGB values. Surprisingly, the noisiness of these visualizations varies substantially for models with similar downstream performance, suggesting a much larger gap in feature quality than indicated by quantitative evaluation. We propose to fix this issue by using an alternative visualization method based on the Laplace operator that inherits all advantages of PCA, while taking the spatial structure of the data into account. We first show that models with noisy PCA feature maps indeed contain linear components with substantially less noise, qualitatively explaining the relatively modest performance differences on dense tasks. Interestingly, the method reveals an opposite problem for the DINOv3 model: PCA visualizations overestimate the spatial smoothness of the model's representations by encoding a strong spatial bias. The proposed Laplacian PCA effectively removes this bias, revealing additional semantically meaningful structure. Quantitatively, we demonstrate that for all tested models linear components identified by Laplacian PCA align better with dimensions relevant for semantic segmentation and monocular depth estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.