Beyond Uniform Soft Thinking: Adaptive Gating and Calibration of Soft Tokens for Latent Reasoning
Abstract
Decoding-time soft thinking extends discrete chain-of-thought reasoning by feeding back a probability-weighted mixture of token embeddings rather than a single-token embedding. However, applying such mixtures at every step does not distinguish between steps with a clear leading token and those with several plausible continuations. Recent selective methods decide when to activate soft mixing, but construct mixtures with a fixed candidate budget or over the full vocabulary. Selecting the steps does not determine how widely to mix, since their next-token distributions can still differ in candidate count and probability concentration. We therefore propose Adaptive Gating and Calibration of Soft Tokens (AGC-ST), a training-free decoding method that determines both when and how widely to mix from the local next-token distribution. Its gate combines entropy and the top-1/top-2 probability margin to select steps where probability mass is broadly distributed and the two most likely tokens have similar probabilities. For each gated step, the Rényi-0.5 effective number, which reflects both candidate count and probability concentration, sets a support cap. A cumulative-mass prefix then selects the final mixing support within this cap. We evaluate AGC-ST on five reasoning models from the Qwen and Llama families, spanning 7B–32B parameters, across five benchmarks covering mathematical, coding, and scientific reasoning. With one shared hyperparameter configuration, AGC-ST achieves the highest accuracy among the evaluated methods in all 25 model–benchmark settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.