Quotient-First Decoding: Sharp Truncation Gaps Without Support Loss
Abstract
Can a truncation decoder change its output law when only the representation of probability mass changes? We analyze exact mass-preserving refinements relative to a supplied outcome quotient. We characterize exactly the outcome laws reachable by nucleus sampling without changing its support. At each fixed parent law, the fine-uniform-refinement limit simultaneously maximizes every convex f-divergence from the reference decoder. The sharp one-step total-variation (TV) supremum is ; its horizon- counterpart is . We identify the mechanism under threshold-prefix compatibility: refinement thins the boundary outcome, preserves the conditional law above it, and can strictly sharpen the distribution. A binary likelihood-ratio factorization explains how this distortion propagates along autoregressive paths. Quotient-first decoding removes it by aggregating before truncation; classical conditional-preserving projection characterizes its unique faithful lift. In state-preserving experiments on three real language models and 120 held-out ARC-Challenge, GSM8K, and MMLU prompt families, Qwen3-8B reaches path TV (95% confidence interval (CI) [0.176, 0.195]). Uniform refinement reaches 0.256 ([0.245, 0.267]) with identical path support. An exact, locally canonical byte quotient gives ; quotient-first replay agrees with the reference kernel to . These results explain representation dependence and establish an exact, conditional-faithful remedy for the supplied outcome quotient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.