SoftSphere: Understanding and Mitigating Collapse in Latent Reasoning
Abstract
Latent reasoning feeds continuous representations back into a language model, but repeated feedback can reduce their diversity. We study and characterize this representation collapse. Our analysis shows how contracting feedback suppresses trajectory differences and how nondegenerate token sampling can replenish variation. Inspired by these observations, we propose SoftSphere, which projects each soft token onto a sphere whose radius matches the mean vocabulary embedding norm. This correction preserves direction and introduces no trainable parameters or auxiliary loss. We establish guarantees for preserving temporal effective rank and slowing state alignment. We evaluate SoftSphere on both token-anchored and noise-anchored latent reasoning modes across a variety of RL tasks, which show consistent improvements in both final reward and diversity. SoftSphere also induces fewer repetitions in the reasoning chain and greater trajectory effective rank. These findings support controlling geometry as a simple way to improve the usefulness of RL rollouts in latent reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.