Adaptive Soft Thinking for Controlling Candidate Influence Across Reasoning States
Abstract
Chain-of-Thought (CoT) reasoning in large language models (LLMs) proceeds through discrete token generation, requiring a single token to be selected at each step even when multiple alternatives remain plausible. Latent reasoning addresses this limitation through continuous intermediate representations, and Soft Thinking provides a training-free approach by constructing Concept Tokens as probability-weighted mixtures of multiple token candidates. However, our preliminary analysis shows that the effect of candidate integration varies across reasoning states and is partially reflected in the current state, motivating state-dependent control of Soft Thinking. Accordingly, we propose ST-Scheduler, a lightweight Soft Thinking scheduler that controls candidate influence through state-dependent Concept Width selection while keeping the underlying LLM frozen. We optimize ST-Scheduler with reinforcement learning with verifiable rewards, using final reasoning outcomes to learn Concept Width decisions throughout reasoning. Across multiple reasoning models and mathematical and scientific reasoning benchmarks, ST-Scheduler outperforms standard CoT, fixed-width Soft Thinking, and existing training-free Soft Thinking methods, while achieving higher solution coverage with fewer sampled trajectories. Code is available at https://anonymous.4open.science/r/ST-Scheduler-FC3F.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.