ToolHallu: Mitigating Hallucination in Agentic Tool Calls via Reinforcement Learning
Abstract
Large language model (LLM)-based agents have been applied across various tasks. A key factor underpinning their powerful capabilities is the use of external tools. While LLMs provide the basic understanding required for tool invocation, their susceptibility to hallucination continues to threaten the reliability of LLM-based agents, making hallucination mitigation imperative for reliable agent deployment. In this work, we propose a novel reinforcement learning (RL) framework named ToolHallu to tackle tool-call hallucination. We first systematically analyze hallucination patterns specific to agentic tool invocation. Based on this, we decompose tool-call hallucinations into distinct error types and assign differentiated rewards to provide interpretable training signals. To further address the varying difficulty of these errors, we introduce a curriculum learning schedule to adjust the optimization process, which first guides the model to learn accurate tool selection and then progressively shifts the optimization focus towards more fine-grained parameter-level hallucinations. Extensive experiments across multiple models and benchmark datasets demonstrate that our method consistently reduces hallucination rates and improves invocation reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.