Sensitivity or Criterion? Measuring and Moving When Language Models Act
Abstract
Before an agent acts it must decide whether any action is warranted. Benchmarks that ask this question report one accuracy, which cannot distinguish a model that fails to recognize a warranted action from one that recognizes it but is reluctant to act. Signal-detection theory separates the two as sensitivity and criterion, but a criterion estimated from authored binary labels depends on each benchmark's mix of examples and cannot be compared across benchmarks. We introduce FlipAct, a diagnostic built on a qualitative representation of what tools do to physical states. A rule over annotated tool properties computes each question's answer and a graded evidence score, with the Yes/No boundary fixed at zero for every model, so every answer is computed, not written by a model. Questions are built from 2x2 quads that hold the required knowledge fixed while the state flips. The slope of a model's Yes-rate against the score is its sensitivity, and the score where it crosses one half is its criterion. Across 44 open models, sensitivity rises with size within a family, and its ranking agrees with three of four external when-to-act benchmarks. Models respond to the sign of the evidence rather than its size, accepting actions that leave the state no better about as often as helpful ones. When Yes means act, most models are conservative and stay so when the question is rephrased. In some models this is a preference for answering No, and the criterion is not set by size. Activation steering moves the criterion with the weights fixed, and distilling a model from its own steered copy moves it in the weights, without a significant change in sensitivity. The distilled shift carries to six external benchmarks with sensitivity kept, including two without a Yes/No answer. Accuracy rises where the model started conservative and falls where it started permissive.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.