acceptodds
Under review as a conference paper at ICLR 2027

Which Steps Matter? Graph-Guided Selection for Single-Step Reinforcement Learning

Abstract

Reinforcement learning for long-horizon agents often requires generating complete interaction trajectories before obtaining a training signal. Repeated reasoning, tool calls, and environment responses make this process expensive, even when the decision of interest occurs much earlier. We propose KStep, a framework that concentrates policy optimization at structurally important decision contexts. From existing trajectories, KStep constructs local state-transition graphs and scores intermediate landmarks by the relative loss of connectivity after their removal. Selected steps become single-step reinforcement learning tasks: the policy receives the full recorded prefix, samples a group of next-turn responses, and updates using group-relative advantages from a reference-action matching reward. Unlike ATLaS, which applies supervised fine-tuning to selected steps, this procedure optimizes newly sampled actions without completing the remaining trajectory. The framework accommodates mathematical reasoning and ALFWorld through domain-specific state adapters. On AIME, KStep exceeds an ATLaS-selection-plus-single-step-RL baseline on two backbones by 3.06 and 1.60 macro-average percentage points. On ALFWorld’s seen split, the corresponding success-rate gains are 6.43 and 3.58 points. Full-rollout GRPO achieves higher final-task performance on both domains. The resulting formulation separates where an agent practices from how its sampled decisions are optimized.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.