LaDex: Latent State Divergence Incentivizes Self-Exploration in LLM Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for enhancing the reasoning abilities of large language models (LLMs). However, vanilla group-based RLVR relies on sparse final-answer rewards, offering limited exploration guidance. Existing exploration algorithms primarily focus on token-level signals, which may reward superficial variation rather than guide the discovery of genuinely diverse and promising reasoning paths. With the observation that trajectories following different reasoning paths show greater divergence in latent space, we propose LaDex, a novel RL method that incentivizes LLM self-exploration through latent-space divergence. For a given prompt, LaDex first generates a group of reasoning trajectories, then segments each full reasoning trajectory into discrete stages and extracts their latent states. Within each reasoning stage, LaDex computes latent divergence by comparing latent states from different trajectories in the group. To balance exploration and exploitation, LaDex introduces a Latent Classifier Network (LCN) that monitors meaningful exploration within each reasoning stage and adjusts the latent-divergence signals accordingly. These signals are then aggregated into intrinsic rewards to guide policy optimization toward effective self-exploration. Experiments on diverse reasoning benchmarks demonstrate that LaDex significantly improves exploration efficiency and consistently outperforms strong baselines. Our code is available at \url{https://anonymous.4open.science/r/LLMRL-CDE2/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.