acceptodds
Under review as a conference paper at ICLR 2027

RaTI: Reliability-aware Truth Inference for Multi-Model Question Answering

Abstract

Large language models (LLMs) have demonstrated strong reasoning capabilities on question answering (QA), while factual errors remain common due to limitations in knowledge coverage. Combining responses from multiple LLMs can improve answer reliability through their complementary knowledge. A central challenge in multi-model QA is determining how much each candidate should influence the final answer because factual reliability varies across candidates. Existing methods typically use an additional LLM for aggregation or treat agreement among generated answers as confidence. However, response agreement does not necessarily indicate factual reliability. Several models may generate the same incorrect answer, giving an unreliable candidate strong consensus and suppressing a less frequent correct answer. To address this limitation, we propose RaTI (**R**eliability-**a**ware **T**ruth **I**nference), a novel reliability-aware truth inference framework for robust multi-model QA aggregation. Rather than relying solely on response consistency, RaTI models the factual reliability of each candidate through the internal representations of LLMs. It further converts the estimated hallucination risks into credibility weights to calibrate semantic Minimum Bayes Risk inference. In this manner, RaTI enables high-reliability candidates to dominate the final aggregation while suppressing misleading false consensus, thereby reducing the reliance on response agreement signals in existing methods. Experiments on six QA benchmarks show that RaTI consistently improves over individual models and representative collaboration baselines, while further analyses confirm the effectiveness of reliability estimation and its robustness under increasing candidate error rates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.