TACO: Improving Uncertainty Decomposition in Black-Box LLMs via Target-Aware Clarification Optimization
Abstract
Uncertainty decomposition is essential for trustworthy reasoning with large language models (LLMs). Recent approaches for black-box LLMs uncertainty decomposition leverage input clarifications to construct ensembles over different interpretations of the same input. However, the effectiveness of such methods critically depends on the quality of generated clarification sets. Existing clarification generation approaches typically treat clarification generation as a general text generation task and overlook its target-dependent nature: the same clarification may be interpreted differently by different target LLMs, introducing additional uncertainty and degrading uncertainty decomposition. To address this limitation, we propose TACO, a Target-Aware Clarification Optimization framework for black-box LLM uncertainty decomposition. TACO formulates clarification generation as a reinforcement learning problem and directly optimizes the clarification generator using feedback from the target LLM through two complementary reward mechanisms: an uncertainty decomposition reward that improves the separation between ambiguous and clear inputs, and an entropy alignment reward that encourages ambiguity-aware diversity while preserving consistency for clear inputs. Extensive experiments demonstrate that TACO consistently improves uncertainty decomposition performance across diverse target LLMs and tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.