UniMAC: Unified Multi-modal Auxiliary Learned Image Compression
Abstract
Learned image compression methods typically operate solely on RGB data, limiting their ability to benefit from complementary information provided by other sensing modalities. Although recent approaches have explored auxiliary inputs such as infrared images, depth maps, point clouds, and stereo views, they generally require a dedicated architecture for each modality, making them difficult to scale and deploy across heterogeneous sensing scenarios. To address this limitation, we propose UniMAC (Unified Multi-modal Auxiliary Learned Image Compression), a unified framework that enables a single compression model to leverage diverse auxiliary modalities. We first employ modality-specific input heads to account for the distinct characteristics of heterogeneous sensor data and project them into a unified feature space. The resulting features are subsequently processed by a shared cross-sensor auxiliary feature extractor, which distills modality-agnostic representations relevant to image compression. These representations are then integrated into the primary RGB compression pipeline to improve its rate–distortion performance. Experiments across diverse cross-sensor settings demonstrate that UniMAC effectively exploits complementary auxiliary cues while avoiding the need to design or train a separate compression model for each modality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.