MS-Point: Multi-Modal Hierarchical Spectral Learning for Point Cloud Understanding
Abstract
Object-level point–image alignment leaves a structural observability gap: agreement between pooled descriptors does not directly constrain intermediate representations or distinguish neighborhood arrangements that produce the same readout. We introduce MS-POINT, a cross-modal pretraining framework that makes internal structure an explicit target of representation learning. At selected encoder stages, complementary neighborhood-consensus and deviation responses expose distinct structural statistics before pooling. Branch-specific visual targets, crosslevel coordination, and adaptive fusion jointly preserve shared object semantics and encourage useful specialization. A common interface of intermediate features and spatial neighborhoods supports both graph encoders and point Transformers. Evaluations across object classification, few-shot adaptation, dense prediction, and low-label scene learning demonstrate the transfer value of this design. Same-backbone interventions and capacity- and computation-matched controls connect these improvements to the structure of supervision. Together, our findings establish supervision of intermediate structural states as a concrete design principle for learning transferable, geometry-aware representations across modalities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.