CALIBRATING LLM EXPLANATIONS TO HUMAN DECISIONS WITH REINFORCEMENT LEARNING
Abstract
Explanations of LLM outputs are considered an important way to make model behavior legible, helping human users make good decisions on downstream tasks. However, default training approaches for improving explanations target model task performance or human preference rather than human decision quality. To bridge this gap, we introduce HD-RL, a reinforcement learning framework for optimizing LLM explanations for human decision making. Our approach calibrates a human decision simulator on observed human decisions, applying a proxy reward model to optimize the LLM to generate explanations that maximize human decision-makers' utility. We apply this framework to AI-generated text detection. While state-of-the-art closed-source LLMs can perform this detection autonomously with an average accuracy of 76.4%, the explanations they generate for users yield near-random human accuracy (51.3%). By applying HD-RL framework to the open-source Qwen 3.6 model, we generate optimized explanations that raise human detection accuracy to over 65.6%, outperforming RLVR, RLAIF, and other baselines by over 10%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.