acceptodds
Under review as a conference paper at ICLR 2027

Nudgeability: Measuring How Confidence Signals Steer Tool Delegation

Abstract

Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any proposal for such self-reflection mechanism has to answer three questions: *where does the reflective signal come from* (verbal reports, output distributions, hidden states, a separate predictor), *how it is presented to the model* (numerical prediction, confidence token, prompt injection), and *whether it changes the model's subsequent action*. We isolate and study the third question. We intervene on otherwise identical reasoning trajectories by inserting, at a fixed point after the same prompt and reasoning prefix, a single first-person sentence expressing either confidence or doubt. The model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation while holding the preceding reasoning trajectory fixed. We define this change in behavioral response as Nudgeability. It is a property of large language models, measured along two dimensions: *sensitivity*, how strongly confidence and doubt change delegation rates, and *targeting*, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive to the intervention: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points. The larger provider-served models exhibit swings of 53 to 70 percentage points. This responsiveness, however, is poorly targeted. Although a median 42% of induced behavioral flips are well-targeted, this represents only a +2 percentage-point median lift over a random-selection baseline. Thus confidence language provides a strong control surface for delegation, but current models seem to use such signals only weakly in accordance with their actual unaided competence. Nudgeability gives us a simple post-training-free way to evaluate both, sensitivity and targeting, as endogenous self-reflection mechanisms mature.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.