OMNI-TRUTH: A Unified Framework for LLM Hallucination Mitigation
Abstract
Large language models must remain faithful to supplied evidence in closed-domain generation and answer reliably from parametric knowledge in open-domain generation without external evidence. Achieving both behaviors in a single model remains challenging, as improving one can compromise the other. We introduce Omni-Truth, a unified post-training framework for closed-domain faithfulness and open-domain factuality. Omni-Truth first develops a shared reflection-and-correction policy through supervised fine-tuning, then applies closed-domain followed by open-domain reinforcement learning to acquire complementary capabilities. Task-specific multi-teacher on-policy distillation consolidates these capabilities into a single model. Across six benchmarks, Omni-Truth achieves the highest macro-average score among the evaluated Qwen3-8B-based methods (67.54), while balancing closed-domain faithfulness and open-domain factuality. On the closed-domain benchmarks, Omni-Truth achieves an average score of 83.75, only 2.69 points below DeepSeek V4 Flash under our evaluation protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.