From Textbooks to Proactive Agentic Systems: Scaling Long-Horizon Planning via Automated Agentic System Construction
Abstract
Large language models (LLMs) excel at generating locally coherent responses but remain fundamentally , lacking the ability to plan and steer interactions over long horizons. Existing approaches attempt to induce proactivity through prompting or reinforcement learning with verifiable rewards (RLVR), which we show degrade on long-horizon dialogue tasks. A natural solution is to augment LLMs with planners, memory, and user simulators. However, specifying these components and their interactions requires substantial expert design effort, thereby limiting scalability. We address this challenge by introducing (**T**extbooks t**O** **P**roactive **A**gentic **S**ystems), a for automatically constructing executable agentic systems from domain textbooks. Given unstructured instructional textbooks, jointly derives a hierarchical planner over structured decision space (states, actions, and rewards), together with a user simulator and a graph-based memory structure, thereby defining a complete interaction loop for training and evaluation. It further generates structured supervision via uncertainty-aware self-annotation and learns hierarchical policies over the extracted structure using reinforcement learning. Experiments on psychotherapy and persuasive dialogue benchmarks show that achieves state-of-the-art performance across multiple metrics and enables planning over longer horizons, outperforming prompting-based, RLVR and manually designed planning approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.