MambaX‑AD: Region‑Aware Contrastive Learning with Mamba-CNN Cross‑Modulation for Multimodal Anomaly Detection
Abstract
In industrial manufacturing, multimodal anomaly detection combining 2D images and 3D point clouds provides complementary appearance and geometric cues for defect inspection. However, existing unsupervised cross-modal mapping methods typically rely on passive fitting of normal correspondences without explicit negative constraints, which can leave an overly broad normal feature boundary and weak discriminative sensitivity. Moreover, static cross-modal mappings have limited capacity to reconcile the asymmetric local details and global context of the two modalities, resulting in a cross-modal alignment bottleneck. To address these two limitations, we propose MambaX-AD. First, we introduce region-aware contrastive learning (RA-CL), a region-aware, mask-conditioned contrastive strategy that uses the spatial mask of feature-level pseudo-anomalies to impose stronger repulsion on corrupted regions while applying a mild balancing penalty to uncorrupted nominal regions. Second, we introduce reciprocal Mamba-CNN cross-modulation mechanism, in which the local CNN and global Mamba branches generate modulation signals for each other and direction- and depth-aware embeddings condition these interactions. Experiments on MVTec 3D-AD achieve the highest mean image-level I-AUROC among the compared methods while maintaining competitive localization performance, and experiments on Eyecandies achieve the highest results on I-AUROC, P-AUROC, and AUPRO@1% among the compared methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.