Conditional Confidence Calibration for Foundation Model Decisions
Abstract
Foundation models increasingly serve as decision components, where discrete outputs require reliable confidence estimates. In safety applications, vision-language and language models make fine-grained judgments about image or text content. These multi-label tasks decompose naturally into binary decisions, each requiring a calibrated confidence estimate. The same need arises whenever a model score is used to accept, reject, or route an output. Standard post-hoc calibration typically learns a global or category-level mapping from model scores to confidence, potentially missing systematic reliability differences across input contexts. We propose conditional confidence calibration, which estimates the probability of a defined binary outcome from the model score, the decision being calibrated, and context available at decision time. Standard global and per-category calibration use only the score and, optionally, the decision category. Our adaptive method additionally fits a contextual calibrator using input descriptors such as text load, visual realism, or request domain, then learns from calibration data how much it should modify the better-supported decision-level estimate. The learned combination can default to that estimate when context is unhelpful, without updating the base foundation model. We evaluate across seven image, image-text, and text safety datasets and three non-safety datasets using Gemma 4 models, with E4B throughout and matched 12B evaluations on two image-text datasets. Across 14 safety comparisons, adaptive calibration improves Macro balanced accuracy over per-category confidence in 12, with mean relative gains of 5.3% on image and image-text tasks and 5.7% on text tasks, and lowers Brier score by 6.3% on average. On the non-safety datasets, it lowers Brier score by 3.2–12.1% and improves balanced accuracy by 2.0–2.8%. The gains also hold across model families: with Qwen3-VL and InternVL3.5 at 4B and 8B, adaptive calibration improves both metrics in all eight comparisons. These results demonstrate that adaptive combination improves both decision and probability quality over non-contextual confidence baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.