acceptodds
Under review as a conference paper at ICLR 2027

Safety after Sharpening: Certified Risk Bounds for Autoregressive Language Models

Abstract

Global power sampling makes a language model's more likely responses more prominent by raising the probability of each complete response to a power and then renormalizing. Recent work uses this approach to improve reasoning without further training karan2026,ji2026, but it can also make incorrect responses more likely. We study how to compute guaranteed lower and upper bounds on the probability of a specified failure after this change of distribution. Our method evaluates selected partial responses with rigorous numerical error bounds and bounds the contribution of every unevaluated continuation. We prove that these intervals can be made arbitrarily narrow for autoregressive transformers that retain all earlier tokens, at any fixed generation length within their context limits, although the worst-case cost is exponential. We also bound failure probabilities when generation stops at the first end-of-sequence (EOS) token or a length limit. When verified bounds on EOS probability and the sum of powered non-EOS token probabilities hold after every reachable partial response, we bound the combined weight of all remaining responses without enumerating them. On a 17-state model with probabilities and stopping bounds checked for every state, these bounds reach interval width while avoiding 8–11 iterations of a dynamic-programming baseline, after both methods evaluate the same next-token probabilities. GPT-2 experiments reach width on 22 of 24 two-token tasks and 8 of 24 longer tasks. In an instruction-tuned model, we certify at least a 9.15-percentage-point increase in the probability of an incorrect initial verdict for one eight-token prompt. In a separate two-token control with equally long answer labels A and B, the effect reverses when we swap their meanings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.