Can AI Solve Imaging Problems? Benchmarking Multimodal AI for Image Formation and Reconstruction
Abstract
Vision-language models (VLMs) and image-generation and editing models have demonstrated strong semantic visual capabilities, but whether they can reason about the physics, sensing mechanisms, and inverse problems underlying imaging remains largely unexplored. We introduce ImagingBench, a benchmark of 22 imaging tasks across five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates multimodal models in three complementary settings: Expert, inverse reconstruction with fixed expert guidance; Planner, inverse reconstruction with model-generated planning; and Forward, forward image-formation simulation. Given a physics-based image-formation model, input measurements, and task context, models directly generate the target image without external code execution or task-specific optimization, probing their native multimodal imaging capabilities. We evaluate leading proprietary and open-source models against representative task-specific baselines. Across tasks, general-purpose multimodal models lag specialized methods, especially on physics-intensive problems such as lensless imaging, event reconstruction, and holography. Planning provides only modest, inconsistent gains, while visually plausible outputs often remain poorly faithful to the reference, revealing a gap between semantic competence and physically grounded imaging. ImagingBench provides a unified framework for quantifying this gap and tracking progress toward multimodal models that reason about imaging physics and inverse problems beyond visual appearance alone. The dataset, code and leaderboard will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.