acceptodds
Under review as a conference paper at ICLR 2027

MS-Point: Multi-Modal Hierarchical Spectral Learning for Point Cloud Understanding

Abstract

Object-level point–image alignment leaves a structural observability gap: agreement between pooled descriptors does not directly constrain intermediate representations or distinguish neighborhood arrangements that produce the same readout. We introduce MS-POINT, a cross-modal pretraining framework that makes internal structure an explicit target of representation learning. At selected encoder stages, complementary neighborhood-consensus and deviation responses expose distinct structural statistics before pooling. Branch-specific visual targets, crosslevel coordination, and adaptive fusion jointly preserve shared object semantics and encourage useful specialization. A common interface of intermediate features and spatial neighborhoods supports both graph encoders and point Transformers. Evaluations across object classification, few-shot adaptation, dense prediction, and low-label scene learning demonstrate the transfer value of this design. Same-backbone interventions and capacity- and computation-matched controls connect these improvements to the structure of supervision. Together, our findings establish supervision of intermediate structural states as a concrete design principle for learning transferable, geometry-aware representations across modalities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.