acceptodds
Under review as a conference paper at ICLR 2027

SpecUnify: Learning Unified Spectral Representations for RGB, Multispectral, and Hyperspectral Vision

Abstract

RGB, multispectral, and hyperspectral images provide different spectral observations of the same visual world, yet modern vision models inherit the sensor's channel configuration as the basis for downstream representation. This couples how a scene is measured with how it is represented for a task, encouraging modality-specific projections, fixed spectral reductions, or separate models and limiting the sharing of representations and supervision across spectral modalities. We ask whether these roles can instead be separated: can heterogeneous spectral observations be mapped into a common task representation whose cardinality is learned from downstream supervision rather than prescribed by the sensor? We present SpecUnify, a unified framework for learning spectral representations across RGB, multispectral, and hyperspectral vision. SpecUnify learns a common spectral representation across modalities by lifting RGB observations into the shared space while directly incorporating multispectral and hyperspectral observations. Within this shared representation, it learns the cardinality of the task representation jointly with the downstream objective, rather than fixing it according to the sensor's channel configuration or a manually chosen spectral reduction. The resulting variable-cardinality representation is converted into fixed-dimensional tokens compatible with pretrained vision models, without reconstructing physical hyperspectral measurements or introducing separate cross-modal pretraining. We evaluate SpecUnify on semantic segmentation and visual tracking under within-modality, cross-modal, and joint RGB-hyperspectral training. The results show that different spectral observations can be processed through a common task representation while maintaining competitive performance across modalities and tasks, and that joint training can exploit complementary supervision from heterogeneous spectral observations. These findings suggest a broader principle for spectral vision: sensor dimensionality need not determine task representation, and heterogeneous spectral observations can instead be unified through a task-learned representation whose cardinality is determined by the downstream objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.