Exact Posterior Distillation: Teaching Language Models What To Believe And Whom To Trust
Abstract
Language model agents gather evidence from sources that may be incomplete or misleading. As evidence arrives, they need to revise both their confidence in possible answers and their trust in individual sources. We introduce BayesClue, a synthetic environment in which models infer a hidden answer from reports by reliable, noisy, and adversarial sources. The known generative process lets us compute exact posterior probabilities over answers and source types after each report. Exact Posterior Distillation (EPD) trains models to predict these probabilities. Across eight models from four families, EPD reduces belief error under changes in evidence length, candidate count, and source proportions. On a separate ten-model comparison, final-answer SFT has higher belief error than every belief-training arm, despite remaining competitive on response generation. EPD also supports answer revision beyond the training horizon and improves average performance on external evidence judgments and contextual appropriateness. These results suggest that explicit supervision of uncertainty can complement training on final answers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.