Vision Through Fields: Learning Real-Sim-Real Reasoning for MLLM Understanding of Multiphysics Combustion Dynamics
Abstract
Multimodal large language models (MLLMs) have made notable progress in image and video understanding, yet their ability to reason about physical processes remains limited. To systematically study this problem, we design a comprehensive multiphysics physical reasoning benchmark grounded in complex engineering scenarios, covering combustion dynamics across diverse environments. The benchmark spans a hierarchy from intuitive physical perception to engineering level reasoning. Evaluations on this benchmark reveal a substantial gap between visual perception and physical state understanding in current MLLMs. To bridge this gap, we propose Vision Through Fields (VTF), a framework for learning real-sim-real physical reasoning, where MLLMs first ground visual observations using aligned simulation fields provided at observed timestamps, reason over field states, field evolution, and cross field relations, and finally transfer the resulting understanding back to visual decisions. To make this reasoning process learnable, we decompose it into real-to-sim grounding, simulation field reasoning, and sim-to-real transfer, and train these capabilities with structured supervision derived from physical fields. This learned real-sim-real reasoning raises mean benchmark accuracy from 42.50% to 54.85% with an MLLM backbone. More broadly, our results highlight the value of simulation derived physical fields as reusable supervision for field guided MLLM physical reasoning. All datasets and code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.