Reasoning-Aware Multicalibrated Reliability Predictors for Language Models
Abstract
As people increasingly rely on large language models (LLMs) for information, model providers need to communicate *meaningful* estimates of whether individual responses are *correct* or, more generally, *reliable*. In the context of hallucinations, existing work uses *multicalibration* to make such reliability predictions meaningful across application-relevant groups, such as prompt topics. Topic groups alone, however, can conceal systematic errors: within the same topic and confidence level, responses with different chain-of-thought (CoT) properties may exhibit opposing errors that cancel on average. In this work, we introduce reasoning-aware multicalibration, a framework that supplements application-relevant topic groups with groups derived from observable properties of CoT traces. We highlight that, theoretically, groupings based on CoT are *"internally important"*, i.e., groupings for whom the multicalibration guarantee is not societally crucial as it is for high-risk topics, but whose inclusion in the multicalibration could easily and significantly improve predictive accuracy of the final reliability predictor. Empirically, we apply our framework to hallucination prediction and unsafe-response detection across a variety of benchmarks, and find that adding CoT groups to topic-based calibration reduces Brier score in both tasks. Interestingly, we find that multicalibrating with respect to CoT groups alone (without any topic information) already improves calibration across topics, suggesting that the predictor's errors depend on how the model reasoned, rather than what it was asked about in the prompt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.