acceptodds
Under review as a conference paper at ICLR 2027

STRIDE: Step-Level Trajectory Reasoning Induced Diversity in Exploration

Abstract

Reinforcement learning with verifiable rewards (RLVR) rewards whether a final answer is correct but does not distinguish between different correct reasoning paths. As a result, training can concentrate probability on a small set of solutions, limiting the benefit of repeated sampling. Existing diversity objectives typically score complete responses and assign the same auxiliary credit to every token, even when only a few reasoning steps make a response distinct. We introduce STRIDE, a post-training method that assigns diversity credit at the level of individual reasoning steps. STRIDE represents each parsed step using a frozen reference model and measures its novelty through regularized reconstruction against the steps of the other eligible rollouts and the preceding steps of its own trajectory. Both correct and incorrect rollouts characterize previously explored directions, but auxiliary credit is assigned only to correct responses. We show that the log-transformed step scores exactly decompose a response's marginal contribution to a regularized log-determinant measure of group diversity. STRIDE converts these scores into bounded bonuses and applies them to the corresponding step tokens during GRPO, requiring neither step-level correctness labels nor additional rollouts. Across five mathematical reasoning benchmarks, STRIDE improves macro Pass@32 from 58.6% to 62.4% on Qwen3-1.7B and from 77.8% to 79.9% on Qwen3-4B-Instruct-2507, while maintaining comparable Pass@1. It also increases macro DistinctStrategies@32 from 0.82 to 0.97 and from 0.92 to 0.97, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.