acceptodds
Under review as a conference paper at ICLR 2027

ObsSpec: Accelerating Agentic Reinforcement Learning with Observation Speculation

Abstract

Tool-using agents alternate between model generation and tool execution, waiting for each tool call to complete before starting the next generation. This sequential execution is costly during agentic reinforcement learning (RL), which requires large batches of multi-turn rollouts. We introduce ObsSpec, an observation speculation pipeline for agentic RL that uses a small world model to predict tool outputs while the environment executes, allowing the policy to begin its next generation early. ObsSpec retains speculative work only when the predicted and actual observations match exactly, making rollout collection lossless, while co-training keeps the world model up to date as the policy learns. On terminal-use tasks, predictions from a 0.6B world model exactly match actual observations in 66.4% of attempts and arrive roughly 16 times faster. Using these predictions, ObsSpec reduces agent latency by 9.3% and 14.2% for 8B and 32B policies. Across controlled experiments, we characterize the regimes in which ObsSpec reduces agent latency under shared policy resources. Finally, because ObsSpec overlaps policy generation with tool execution, it is architecturally compatible with existing techniques for accelerating RL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.