Which Steps Matter? Graph-Guided Selection for Single-Step Reinforcement Learning
Abstract
Reinforcement learning for long-horizon agents often requires generating complete interaction trajectories before obtaining a training signal. Repeated reasoning, tool calls, and environment responses make this process expensive, even when the decision of interest occurs much earlier. We propose KStep, a framework that concentrates policy optimization at structurally important decision contexts. From existing trajectories, KStep constructs local state-transition graphs and scores intermediate landmarks by the relative loss of connectivity after their removal. Selected steps become single-step reinforcement learning tasks: the policy receives the full recorded prefix, samples a group of next-turn responses, and updates using group-relative advantages from a reference-action matching reward. Unlike ATLaS, which applies supervised fine-tuning to selected steps, this procedure optimizes newly sampled actions without completing the remaining trajectory. The framework accommodates mathematical reasoning and ALFWorld through domain-specific state adapters. On AIME, KStep exceeds an ATLaS-selection-plus-single-step-RL baseline on two backbones by 3.06 and 1.60 macro-average percentage points. On ALFWorld’s seen split, the corresponding success-rate gains are 6.43 and 3.58 points. Full-rollout GRPO achieves higher final-task performance on both domains. The resulting formulation separates where an agent practices from how its sampled decisions are optimized.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.