Singular Value Perturbation: Unlocking RL Performance via Diverse Latent Reasoning
Abstract
Latent Reasoning Models (LRMs) replace explicit Chain-of-Thought tokens with continuous recurrent computation, yet their deterministic recurrence limits the exploration needed for effective reinforcement learning. Sampling answer tokens alone cannot explore alternative computations within the latent reasoning phase. We introduce Singular-Value Perturbation (SVP), which diversifies latent trajectories by randomly rescaling coefficients in the fixed SVD basis of a weight matrix. Applied to attention Value projections as SVP-V, it perturbs the messages transmitted by attention while preserving the attention weights for a fixed input. Across multiple LRMs and mathematical reasoning benchmarks, SVP-based configurations achieve the highest solution coverage in all evaluated settings, while intermediate LM-head readouts reveal semantic branching around decision-critical steps. Building on this exploration mechanism, SVP-V-GRPO learns from the relative outcomes of perturbed trajectories while updating only the Value parameters. Starting from COCONUT, it improves GSM8K accuracy from 34.1% to 50.3%, achieving state-of-the-art performance across all six benchmarks among GPT2-based continuous LRMs. These results establish structured latent exploration as an effective mechanism for RL post-training of continuous LRMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.