MARVEL: Masked Autoregressive Visual Understanding and Generation for Spectral Remote Sensing
Abstract
Spectral remote sensing observations support both scene understanding and the generation of unobserved spectral responses. Bringing understanding and generation into unified pretraining allows generative supervision to support transferable representation learning. However, more accurate generation does not necessarily yield better features for understanding. To address this challenge, we introduce MARVEL, a masked autoregressive pretraining framework for spectral understanding and generation. A bidirectional encoder integrates all visible observations, and full-context conditional alignment converts the resulting features into conditions for each target wavelength and spatial position. Spectral content autoregression then uses these conditions to generate responses in wavelength order. Each response is supervised against its own target and re-encoded to guide subsequent predictions. Through this reuse, feedback from subsequent predictions also contributes to shared representation learning. MARVEL achieves competitive performance on 14 downstream tasks. Without additional generation training, it also achieves strong results on spectral generation tasks using Sentinel-2 and Landsat-8 data. Controlled experiments further validate the effectiveness of the proposed design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.