Estimating Uncertainty of Omnimodal Large Language Models via External Omnimodal Embeddings
Abstract
Recent work has investigated Omni-modal Large Language Models (OLLMs), models which can take multiple modalities and their combinations as input, as a way forward to have systems really capable of interacting with their environment. Uncertainty quantification (UQ) for OLLMs, however, remains an open problem: existing methods largely require access to model internals, or expensive self-verbalisation approaches. We propose a model-agnostic UQ framework that scores a Chain-of-Thought response entirely through external omni-modal embeddings, measuring the grounding of the reasoning in the multimodal input and its internal coherence. The resulting confidence estimates require no access to logits, hidden states or gradients of the evaluated model, rather relying on the segmented reasoning trace and, for the aggregated score, a labelled calibration set that we show can be as small as 25 examples. On UNO-Bench, with both a closed and an open model, our scores yield small gains over self-consistency (systematic with the closed generator, mixed with the open one) and are competitive with self-verbalisation and internal-state baselines, especially in open-ended, omni-modal questions where previous state-of-the-art baselines struggle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.