acceptodds
Under review as a conference paper at ICLR 2027

Eliciting Action Uncertainty via Reasoning Interruption for Training Calibrated Tool-Using LLMs

Abstract

Reinforcement learning improves the execution accuracy of tool-using LLMs but often leaves them poorly calibrated, causing the models to remain overconfident in incorrect actions and thereby undermining their reliability and trustworthiness. Existing approaches inject a calibration term into the optimization objective using uncertainty estimates aggregated over the entire reasoning-to-action sequence. Such trace-level estimates conflate linguistic fluency with action correctness: because action-relevant information is sparse, the resulting uncertainty signal is diluted by generic reasoning tokens. To recover an action-faithful uncertainty signal, we interrupt the reasoning trace at step boundaries and regenerate tool calls from each prefix. Our analysis shows that the action's log-likelihood, evaluated under such interrupted prefixes, ranks correct tool calls above failures more reliably than its full-trace counterpart. Based on this observation, we propose UARI (Uncertainty-aware Reasoning Interruption), an RL framework that turns this interrupted-action signal into an uncertainty-aware training objective: candidates are grouped by uncertainty-induced difficulty and the same signal supplies a bounded reward-shaping term, jointly optimizing execution correctness and calibration. By targeting an uncertainty signal decoupled from generic reasoning tokens, UARI addresses miscalibration at the specific juncture of action commitment. Empirical evaluations on the BFCL-v4 and Nestful benchmarks demonstrate that UARI consistently outperforms the standard GRPO baseline in both execution accuracy and the calibration peromance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.