acceptodds
Under review as a conference paper at ICLR 2027

SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization

Abstract

The soft-thinking reasoning paradigm has demonstrated superior performance over traditional discrete-token Chain-of-Thought (CoT) reasoning in various scenarios. However, unlike discrete-token CoT reasoning that can be reinforced through advanced reinforcement learning with verifiable rewards (RLVR) techniques such as group relative policy optimization (GRPO), extending the soft-thinking reasoning with such strong techniques remains challenging. This difficulty stems from the complexities of injecting stochasticity into soft-thinking tokens and updating soft-thinking policies accordingly. As a result, previous attempts to combine soft-thinking with RLVR typically underperform their discrete-token RLVR counterparts. To fully unlock the potential of soft-thinking, this paper presents a powerful policy optimization algorithm, SofT-GRPO. It injects the Gumbel noise into token probabilities with Gumbel-Softmax for controllable stochasticity, and leverages the Gumbel reparameterization trick to achieve accurate credit assignment to LLM soft-thinking policies. We conduct experiments over LLMs ranging from 1.5B to 7B parameters, where SofT-GRPO enables LLMs with soft-thinking to slightly outperform discrete-token CoT GRPO on Pass@1 (+0.13% on average accuracy), and brings a substantial uplift on Pass@32 (+2.19% on average).The codes are available at https://anonymous.4open.science/r/SofT-GRPO-master-A512.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.