Empowering Multimodal Large Language Models for RAW Image Understanding
Abstract
In this paper, we enable multimodal large language models (MLLMs) to reason from RAW measurements by addressing a critical blind spot: existing benchmarks evaluate only post-ISP images, in which processing may weaken low-light, high- light, illumination, color, and texture cues. Current evaluations therefore cannot expose this information loss. We introduce RAW Bench, a diagnostic bench- mark comprising 559 questions over 368 paired scenes across five RAW-sensitive dimensions. Each item is RAW-indispensable, requiring physical evidence pre- served in RAW but weakened in the matched sRGB image. To achieve this crite- rion, RAW Bench uses matched RAW/sRGB inputs, hidden fact anchors, and sep- arate scores for pairwise correctness, fact grounding, and reasoning quality, distin- guishing evidence-grounded answers from semantic guesses. We further propose a dual-pathway RAW-aware adaptation framework that preserves the pretrained semantic interface while injecting compact RAW context into projected visual to- kens, yielding consistent same-backbone gains across two MLLM families, partic- ularly in low-light, dark/highlight, and illuminant reasoning. These results demon- strate the value of sensor-domain evidence for multimodal reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.