acceptodds
Under review as a conference paper at ICLR 2027

Translating Hyperspectral Images into Multimodal Semantics that Vision-Language Models can Better Understand

Abstract

Despite capturing unified spatial-spectral representations, hyperspectral foundation models incur prohibitive pre-training costs and require sensor/scene-specific fine-tuning, severely hindering real-world deployment. This raises a fundamental question: "Can we eliminate domain-specific pre-training and adaptation, directly leveraging pre-trained vision-language models (VLMs) for training-free hyperspectral interpretation?" To address this, we propose HyperMAG, a novel framework that rethinks hyperspectral image interpretation through multimodal semantic augmented generation. Our key insight is that pre-trained VLMs, despite having never seen hyperspectral data, can reliably interpret spectral information when effectively translated into their native modalities. Specifically, HyperMAG translates spectral signatures into complementary textual descriptors (band-indexed values and peak–valley trends) and visual strips (encoding magnitude and local contrast) across multi-scale homogeneous regions, enabling off-the-shelf VLMs to extract rich spatial-spectral semantics. Extensive experiments demonstrate that HyperMAG matches or surpasses both task-specific models and foundation models across diverse downstream tasks, including few-shot classification, single-class classification, target detection, anomaly detection, and change detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.