Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
Abstract
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner–executor systems can fail at either stage: the agent may deviate from a structured plan during execution, or it may faithfully execute a plan that is poorly matched to the task or environment. Final task success alone cannot distinguish between these two sources of failure. We therefore ask: can LLM agents be trusted to execute the plans they commit to, and do different tasks and environments benefit from different planning modes? To study this, we introduce a diagnostic framework for the *Plan Declaration–Execution Gap* and propose *Planning-as-Routing*, in which the LLM declares one of four planning modes: *{Predefined, Sequential, Hierarchical, or Search}*. A deterministic router sends the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve the declared planning structure, especially for longer plans: across three benchmarks, only \(22\)–\(45%\) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from plan execution: pattern-specific executors raise task success from \(0.48\) to \(0.92\) on ALFWorld and from \(0.36\) to \(0.44\) on SWE-bench Verified over Plan+ReAct. However, current LLMs do not reliably select the strongest planning mode for a task, while few-shot examples can improve planning-mode selection for some benchmark–model combinations. Overall, reliable agent planning requires both selecting an effective planning mode and executing it with a matching executor: routing substantially closes the execution gap, while selecting the right mode for each task remains open.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.