acceptodds
Under review as a conference paper at ICLR 2027

From Textbooks to Proactive Agentic Systems: Scaling Long-Horizon Planning via Automated Agentic System Construction

Abstract

Large language models (LLMs) excel at generating locally coherent responses but remain fundamentally , lacking the ability to plan and steer interactions over long horizons. Existing approaches attempt to induce proactivity through prompting or reinforcement learning with verifiable rewards (RLVR), which we show degrade on long-horizon dialogue tasks. A natural solution is to augment LLMs with planners, memory, and user simulators. However, specifying these components and their interactions requires substantial expert design effort, thereby limiting scalability. We address this challenge by introducing (**T**extbooks t**O** **P**roactive **A**gentic **S**ystems), a for automatically constructing executable agentic systems from domain textbooks. Given unstructured instructional textbooks, jointly derives a hierarchical planner over structured decision space (states, actions, and rewards), together with a user simulator and a graph-based memory structure, thereby defining a complete interaction loop for training and evaluation. It further generates structured supervision via uncertainty-aware self-annotation and learns hierarchical policies over the extracted structure using reinforcement learning. Experiments on psychotherapy and persuasive dialogue benchmarks show that achieves state-of-the-art performance across multiple metrics and enables planning over longer horizons, outperforming prompting-based, RLVR and manually designed planning approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.