acceptodds
Under review as a conference paper at ICLR 2027

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Abstract

How should an agent decide when and how to plan? A dominant approach builds the agent as a reactive policy with adaptive computation (e.g., chain-of-thought reasoning), trained end-to-end with the expectation that planning will emerge implicitly from sufficient data and compute. Without control over the presence, structure, or horizon of planning, however, these systems typically increase reasoning length dramatically during training, leading to inefficient token consumption that does not reliably translate to accuracy gains. We argue that efficient agentic reasoning benefits from a decomposition of decision-making into three interacting systems: **simulative reasoning** (System II) that grounds deliberation in future-state prediction using a world model, rather than unconstrained chain-of-thought; **self-regulation** (System III) that decides when and how deeply the agent plans at each turn through a learned **configurator**; and **reactive execution** (System I) that handles fine-grained reasoning and action. To test this, we develop SRAM (Self-Regulated Simulative Reasoning Agentic LLM), which realizes the configurator and simulative planning as distinct stages within an LLM's chain-of-thought reasoning, with the LLM itself serving as the world model in language space. We explore two implementations, v0.1-8B and v1.0-30B, both trained via supervised learning followed by reinforcement learning (RL). Across mathematical reasoning, scientific problem-solving, tabular data analysis, and web information seeking, SRAM-v0.1-8B and SRAM-v1.0-30B achieve overall Pass@1 competitive with systems at 120–357B and 685B–1T parameters, respectively, while SRAM-v1.0-30B consumes 25.8–94.9% fewer reasoning tokens than strong agentic LLMs of similar scale. Analysis reveals that RL increases average planning horizon by 22.8% while planning frequency grows only 2.0 percentage points, indicating that the model learns to plan *further ahead* rather than *more often*.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.