RL-Driven Autonomous Hallucination Reduction in High-Precision Agentic Applications
Abstract
Reliable hallucination detection for deployed conversational AI requires both accuracy and real-time latency—constraints existing approaches satisfy only in isolation. Generative LLM judges achieve strong accuracy but require a full autoregressive generation per turn; lightweight probes on uncertainty features offer sub-millisecond inference but degrade on the hardest cases. We identify and formally characterise a previously undocumented failure regime—confident-fluent (CF) hallucination—in which a language model assigns near-zero entropy at the verdict token (0.0003 nats, a 140x compression relative to the entropy of hallucinations that are caught) while producing factually incorrect output. Through controlled ablations on a human-annotated benchmark spanning four failure taxonomies, we show this constitutes a hard feature-space ceiling: we prove that the Youden index of any decision rule measurable in the self-assessment features is bounded above by the total-variation distance between the CF and correct-confident feature distributions, which we measure at . No linear or interaction combination of standard self-assessment features can separate CF hallucinations from true PASS turns, regardless of probe architecture or training-data scale. To address this ceiling we introduce two components and keep them distinct: POLYGRAPH, the probe that scores each turn and serves as the reward model, and VERITAS (Verification via Embedded Representations in Interactive Tuning for Agentic Safety), the autonomous Reinforcement Learning + Reward Model (RL+RM) loop that consumes it to adapt the agent's generation policy without per-turn human labels. The probe supplies a dense per-turn reward to a PPO fine-tuning loop: the agent is penalised for confident-fluent hallucinations and rewarded for grounded responses. Against an 11-method benchmark on 4,600 AI-coaching conversations, the static probe reaches 85.5% accuracy at a 12.3% hallucination rate, surpassing LLM-as-judge baselines by 5.5 points at 100x lower latency. After 500 RL steps the autonomous loop reaches an 8.0% hallucination rate—below the 10% production safety target with zero human labels beyond the initial bootstrap. Code, synthetic training data, and a replication package are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.