Diversity Without Discovery: Why Chain Perturbation Fails to Explore in Latent Test-Time Scaling
Abstract
Parallel test-time scaling (PTTS) improves reasoning by sampling multiple chain-of-thought (CoT) traces and aggregating them, marginalizing over the reasoning paths the model inherently finds plausible. Various SOTA latent-reasoning models instead evolve continuous hidden states under deterministic dynamics, following one trajectory per input with no native randomness to sample from. Recent work therefore perturbs the latent chain at test time and presents the resulting answer diversity as exploration potential in continuous space. To test this, we fairly compare latent perturbation methods with standard explicit reasoning on a shared backbone and find that, for unconstrained latent thoughts, i.e., thoughts not restricted to the embedding space, aggregating perturbed latent chains brings no gain over the no-sampling baseline, whereas explicit reasoning turns the same budget into accuracy gains. Our formal analysis demonstrates that explicit reasoning marginalizes over the trajectory distribution stochastically exhibited by the model, while a perturbed state comes from no such distribution, so its vote carries no belief of the model and no guarantee of favoring the correct answer. Empirically, the added diversity comes predominantly from borderline samples and, unlike with explicit reasoners, a correct answer reached through perturbation carries no gain in confidence. As such, even a frozen latent chain, perturbed at most once at the answer step, can be calibrated to reproduce or outperform every full-chain perturbation method we tested at no additional reasoning-token cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.