acceptodds
Under review as a conference paper at ICLR 2027

The Path Matters: Trajectory-Guided Semi-Supervised Reinforcement Learning for LLM Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs), but its reliance on reference answers limits training when annotations are scarce. Existing approaches exploit unlabeled questions through self-generated rewards, and recent semi-supervised methods use limited labeled data to select reliable training examples. However, answer agreement and sample-level selection provide limited guidance for the quality of individual reasoning trajectories: different rollouts that reach the same pseudo-answer receive identical correctness feedback, even when their reasoning differs in reliability. Motivated by this consideration, we propose , a trajectory-guided semi-supervised reinforcement learning framework for LLM reasoning. maintains a dynamic reasoning graph anchored by trajectories whose answers are verified on labeled examples, and uses related questions to provide auxiliary support for each unlabeled rollout. It separates pseudo-label confidence from trajectory guidance, using the former to weight each prompt's policy update and the latter to shape rollout-level group-relative advantages. A separate admission rule conservatively expands the graph with pseudo-labeled examples while allowing broader participation in training. We design an evaluation on six mathematical reasoning benchmarks under three labeled–unlabeled data allocations to assess reasoning performance and label efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.