acceptodds
Under review as a conference paper at ICLR 2027

Mask Components, Not Bands: Rethinking Masked Autoencoding for Hyperspectral Imagery

Abstract

Hyperspectral imagery (HSI) captures hundreds of contiguous wavelengths per pixel, enabling material identification through spectral differences, but labeled data is scarce. Masked autoencoder (MAE) pretraining has emerged as the dominant self-supervised approach for learning HSI representations from unlabeled data. Existing HSI MAE methods adopt input-space masking from RGB imagery and video, a recipe designed around spatial redundancy. HSI distributes its redundancy differently. Neighboring pixels are often non-redundant, while adjacent wavelengths are so highly correlated that masked bands remain linearly recoverable from visible ones until most of the spectrum is removed. Input-space methods therefore rely on masking ratios of up to 90%, which make reconstruction difficult only by leaving the encoder a sparse view of each pixel's spectrum. We propose CoMask, which exploits this spectral correlation by projecting HSI pixels onto a small number of basis components that span all wavelengths. CoMask masks in this component basis, where zeroing a component removes its contribution from every wavelength. The encoder therefore sees every wavelength at reduced fidelity, with degradation controlled by how many components are masked. Reconstruction difficulty then depends on how much spectral signal is removed rather than on which wavelengths are hidden. That dependency aligns the task with the low-dimensional structure of HSI data. Under linear probing on Houston2018 at ViT-Tiny, CoMask exceeds every input-space masking configuration in our sweep, leading the strongest by 2.7 points of overall accuracy averaged over three pretraining seeds. On EnMAP-DFC the lead appears under finetuning rather than probing, at 2.0 points of macro accuracy over the strongest configuration in that sweep. Pretrained on the foundation-scale Spectral Earth corpus, CoMask outperforms an input-space baseline of the same architecture and recipe on all six downstream segmentation tasks, by 0.9 to 2.6 micro-IoU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.