acceptodds
Under review as a conference paper at ICLR 2027

Dropout-GRPO: Variational Attention Dropout for Continuous Latent Reasoning

Abstract

Large Language Models(LLMs) have achieved considerable success in various reasoning tasks by utilizing Chain-of-Thought(CoT) reasoning with intermediate language tokens. An emerging family of models utilize the latent continuous space of the LLMs in pursuit of robustness and efficiency, bypassing the restrictions of linguistic fluency and verbosity. However, often such models are deterministic by design and require complex architectural modifications to achieve variation in the reasoning trajectory necessary for critic-free post-training methodologies such as GRPO. We investigate dropout as a lightweight mechanism for exploring alternative latent trajectories without redesigning the recurrence. By applying dropout to attention probabilities, we perturb the continuous latent trajectories. We use rollout-specific attention-mask patterns shared across latent steps and replay the corresponding dropout realization during policy updates. This supplies rollout diversity without changing the latent recurrence architecture. We evaluate this approach on arithmetic benchmarks and support it with a mathematical analysis of replay and the scope of the resulting policy update. The results validate that latent reasoning models can benefit from post-training through reinforcement learning with verifiable rewards by using dropout. The corresponding code is shared via supplementary materials, and the repository will be made public upon acceptance of the paper.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.