acceptodds
Under review as a conference paper at ICLR 2027

Empowering Multimodal Large Language Models for RAW Image Understanding

Abstract

In this paper, we enable multimodal large language models (MLLMs) to reason from RAW measurements by addressing a critical blind spot: existing benchmarks evaluate only post-ISP images, in which processing may weaken low-light, high- light, illumination, color, and texture cues. Current evaluations therefore cannot expose this information loss. We introduce RAW Bench, a diagnostic bench- mark comprising 559 questions over 368 paired scenes across five RAW-sensitive dimensions. Each item is RAW-indispensable, requiring physical evidence pre- served in RAW but weakened in the matched sRGB image. To achieve this crite- rion, RAW Bench uses matched RAW/sRGB inputs, hidden fact anchors, and sep- arate scores for pairwise correctness, fact grounding, and reasoning quality, distin- guishing evidence-grounded answers from semantic guesses. We further propose a dual-pathway RAW-aware adaptation framework that preserves the pretrained semantic interface while injecting compact RAW context into projected visual to- kens, yielding consistent same-backbone gains across two MLLM families, partic- ularly in low-light, dark/highlight, and illuminant reasoning. These results demon- strate the value of sensor-domain evidence for multimodal reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.