acceptodds
Under review as a conference paper at ICLR 2027

M-Plan: Comprehensive Planning For LLM Agents with Memory Arrangement

Abstract

Large language model (LLM)-based agents have recently demonstrated outstanding performance in single-turn tasks, yet they continue to struggle with long-horizon tasks, particularly in environments without prior training. A key limitation of single LLMs is their difficulty in processing long contexts, which stems from weaknesses in continuous decision-making and reasoning capabilities. Although existing multi-agent systems (MAS) equipped with long-term memory bases can partially address this gap, they mostly rely on knowledge reuse or summarization—an approach that is often constrained by the specific requirements of the current task. In contrast, MAS characterized by task-specific guidance often lack comprehensive long-term memory designs. To address these challenges, we propose M-Plan, a closed-loop planning framework that integrates long-term and short-term memory. By combining global-local and past-current guidance, M-Plan enhances agents' performance on multi-turn tasks while maintaining a plug-and-play design. During task execution, agents receive task-specific milestones and insights, and are provided with local hints only when confusion arises. These guidance generation and utilization construct the short-term memory. After task completion, relevant information about the current task is optimized or reorganized before being sent to the long-term memory storage for future retrieval. Extensive experiments on multi‑turn tasks including ALFWorld, WebShop, and PDDL across diverse models demonstrate that M‑Plan outperforms multiple baseline methods in both accuracy and efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.