acceptodds
Under review as a conference paper at ICLR 2027

Reinforcement Learning Done Right for Explicit Confidence Targets

Abstract

Large language models (LLMs) can produce plausible but incorrect answers when they are uncertain, which can mislead users and undermine model reliability. Training models to abstain with reinforcement learning is a natural approach, but most existing methods use a fixed error penalty and therefore cannot adapt their behavior to different confidence requirements at inference time. We study explicit confidence targets, where a user specifies a target p in the prompt and the model directly chooses whether to answer or abstain, with the goal of achieving an accuracy of at least p among answered questions. However, learning to respond to different confidence targets with a single model has proven difficult, and prior attempts have failed to distinguish different targets. We show that this failure comes from two problems: the reward signal that distinguishes different values of p is often lost during optimization, while the original error penalty grows without bound as p approaches one. We propose RLEC (Reinforcement Learning with Explicit Confidence Targets), which preserves the learning signal needed to distinguish confidence targets, introduces a stable reward family that keeps the desired answer condition unchanged, and uses reward scheduling to balance effective learning and training stability. Together, these designs enable the first successful training of a single LLM for explicit confidence target control. Extensive experiments across diverse benchmarks and model scales show that RLEC largely preserves general capabilities and achieves competitive calibration performance. RLEC also supports fine-grained target control without an external decision system or threshold calibration after training. By making explicit confidence targets part of the trained policy, RLEC opens a new and more direct path toward behavioral calibration. Our code is available at: [https://anonymous.4open.science/r/RLEC-ICLR-48E5](https://anonymous.4open.science/r/RLEC-ICLR-48E5).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.