MIGE: Multi-Level Influence-Guided Exploration in Reinforcement Learning for LLM Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) substantially improves reasoning in large language models, but policies can collapse to a small set of familiar solution paths. Existing exploration methods mainly preserve token entropy or reward trajectory diversity. Yet high token uncertainty need not mark a consequential reasoning choice, and textually diverse trajectories may follow the same underlying strategy. To address this limitation, we propose Multi-Level Influence-Guided Exploration (MIGE), which guides exploration according to how decisions affect subsequent reasoning. At the strategy level, MIGE computes the log-likelihood of complete executions under alternative strategy contexts to reward strategies with distinct downstream effects. At the step level, it resamples continuations at candidate token positions and rewards forks that change the next intermediate conclusion. Across six mathematical and scientific reasoning benchmarks and two model scales, MIGE achieves the best average performance among standard RLVR and representative exploration baselines while largely preserving successful-solution coverage at Pass@. Controlled analyses show that influence scores distinguish strategy changes from surface rewrites and locate reasoning forks more reliably than output-distance and uncertainty baselines, supporting downstream influence as an effective signal for meaningful exploration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.