MuCCA: Multilevel Human Color Representation-guided Complementary Alignment for Neural Visual Decoding
Abstract
Neural visual decoding links brain activity to visual content and supports the development of brain–computer interfaces. However, existing approaches to visual supervision adaptation do not explicitly account for the multilevel organization of color processing in human vision, potentially limiting neural–visual alignment. To address this limitation, we propose Multilevel Human Color Representation-guided Complementary Alignment (MuCCA) for visual decoding. MuCCA constructs RGB, LMS, DKL-inspired, and CIELUV views of each stimulus to provide supervision informed by cone responses, opponent relationships, and perceptual color attributes. Four independently trained Residual Projection Encoders align neural recordings with frozen CLIP ResNet-50 features of these views through symmetric contrastive learning. A separate original-image branch jointly trains an ATM neural encoder and a ViT-H-14 image encoder for generation. For retrieval and classification, the four RPE branches produce individual rankings, and their candidate union quantifies complementary decoding successes. For reconstruction, a diffusion prior predicts the primary condition from the ATM embedding, while three linear projections map the LMS-, DKL-, and CIELUV-supervised neural embeddings into the same conditioning space. Feature supervision and cross-space consistency regularize these projections. Four IP-Adapters integrate these conditions into a pretrained SDXL generator for reconstruction. Extensive experiments on THINGS-EEG and THINGS-MEG demonstrate that MuCCA outperforms state-of-the-art methods across retrieval, classification, and reconstruction. The code is available at https://anonymous.4open.science/r/MuCCA-8471-842D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.