acceptodds
Under review as a conference paper at ICLR 2027

TraceConf: Visually Traceable and Calibrated Confidence in Multimodal Large Language Models

Abstract

As multimodal large language models are increasingly used, confidence estimates are becoming central to assessing the reliability of their predictions. Existing calibration objectives, however, primarily ask whether an answer is likely to be correct, without revealing which visual evidence warrants that confidence. This limitation arises because binary correctness assigns the same calibration target to predictions with different evidential bases, allowing confidence to be statistically calibrated while remaining visually untraceable. We introduce TraceConf, a post-training approach that jointly generates declared visual evidence, an answer, and a confidence choice. TraceConf counterfactually assesses the necessity and sufficiency of the declared regions by removing and isolating them. The resulting evidence gate rewards evidence-grounded response generation and weights a local proper confidence loss, while the confidence-choice token is excluded from the shared sequence-level GRPO advantage to prevent response-level rewards from misassigning credit to confidence. Because the gate is determined before the confidence report, evidence weighting preserves the proper conditional target wherever the gate is positive. Across six multimodal benchmarks, TraceConf improves confidence calibration on average and error discrimination across all benchmarks while maintaining competitive answer accuracy. Further experiments establish a direct link between model confidence and its declared visual evidence: removing the declared regions causes larger confidence drops, while retaining only these regions better preserves confidence than retaining their complementary regions. Together, these results demonstrate that visually traceable confidence can enhance multimodal reliability without compromising answer quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.