acceptodds
Under review as a conference paper at ICLR 2027

Hallucination-Corrected Action Compression via LLM-RL Coordination for Task-Oriented Dialogue policy optimization

Abstract

Task-oriented dialogue policy optimization remains challenging due to the large and highly structured action space. Large language models (LLMs) can provide state-conditioned action priors to compress this space, but their hallucinations may introduce task-irrelevant, constraint-violating, or infeasible actions. Such unreliable actions not only undermine action-space compression but also bias downstream action-value estimation and policy learning. To address this issue, we propose Hallucination-corrected Action Compression (HAC), a two-stage framework that corrects the LLM prior before using it to guide dialogue policy optimization. Stage I performs hallucination-induced correction on the original LLM prior. Specifically, an offset LLM is induced to amplify hallucination tendencies, and its predictions are contrasted with those of the original LLM to suppress unreliable preferences. This yields a factually corrected action prior, from which a compressed action set is constructed—producing a more trustworthy action space than directly using the original LLM. Stage II then performs stable RL-based policy optimization solely over this compressed set. A distributional Q-network estimates return distributions for each action, and a categorical cross-entropy objective aligns predicted and target distributions, enabling stable and discriminative learning within a reduced and reliable action space. Experiments on multiple task-oriented dialogue and factuality benchmarks show that HAC achieves superior task performance, improved factual reliability, and strong generalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.