acceptodds
Under review as a conference paper at ICLR 2027

Conditional Confidence Calibration for Foundation Model Decisions

Abstract

Foundation models increasingly serve as decision components, where discrete outputs require reliable confidence estimates. In safety applications, vision-language and language models make fine-grained judgments about image or text content. These multi-label tasks decompose naturally into binary decisions, each requiring a calibrated confidence estimate. The same need arises whenever a model score is used to accept, reject, or route an output. Standard post-hoc calibration typically learns a global or category-level mapping from model scores to confidence, potentially missing systematic reliability differences across input contexts. We propose conditional confidence calibration, which estimates the probability of a defined binary outcome from the model score, the decision being calibrated, and context available at decision time. Standard global and per-category calibration use only the score and, optionally, the decision category. Our adaptive method additionally fits a contextual calibrator using input descriptors such as text load, visual realism, or request domain, then learns from calibration data how much it should modify the better-supported decision-level estimate. The learned combination can default to that estimate when context is unhelpful, without updating the base foundation model. We evaluate across seven image, image-text, and text safety datasets and three non-safety datasets using Gemma 4 models, with E4B throughout and matched 12B evaluations on two image-text datasets. Across 14 safety comparisons, adaptive calibration improves Macro balanced accuracy over per-category confidence in 12, with mean relative gains of 5.3% on image and image-text tasks and 5.7% on text tasks, and lowers Brier score by 6.3% on average. On the non-safety datasets, it lowers Brier score by 3.2–12.1% and improves balanced accuracy by 2.0–2.8%. The gains also hold across model families: with Qwen3-VL and InternVL3.5 at 4B and 8B, adaptive calibration improves both metrics in all eight comparisons. These results demonstrate that adaptive combination improves both decision and probability quality over non-contextual confidence baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.