OTU: Process-aware Uncertainty Estimation for Multimodal Reasoning
Abstract
Multimodal large language models (MLLMs) have demonstrated strong capabilities in visual understanding and multi-step reasoning, yet their predictions often lack reliable uncertainty estimates. Existing uncertainty quantification (UQ) methods predominantly operate at the answer level, treating the final response as a whole rather than modeling its internal reasoning process. As a result, they often fail to capture uncertainties caused by ambiguous evidence, imprecise grounding, or errors in intermediate reasoning steps. This limitation is particularly critical in multimodal settings, where uncertainty may stem from both perception and reasoning. In this work, we propose OTU, a process-aware uncertainty estimation framework that models reasoning trajectories as chain graphs and quantifies uncertainty via cross-trajectory consistency. In a fully black-box setting, OTU employs optimal transport to align multiple sampled trajectories and compute a consensus barycenter that captures dominant semantic and structural patterns. Uncertainty is then derived from the deviation of individual trajectories from this consensus, enabling both answer-level confidence estimation and fine-grained step-level reliability analysis. Extensive experiments on multiple multimodal benchmarks show that OTU consistently improves calibration over existing black-box baselines. Moreover, it provides interpretable insights into where and how reasoning fails, highlighting that uncertainty in MLLMs is inherently a process-level phenomenon rather than solely an outcome-level signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.