acceptodds
Under review as a conference paper at ICLR 2027

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

Abstract

People state planning goals in natural language, while classical planners require formal PDDL. Recent agentic frameworks bridge this gap with a verifier-checked refinement loop in which an orchestrator picks, at every step, which repair agent (an LLM-prompted or deterministic PDDL and plan-repair routine) to run next; that orchestrator is itself a prompted frontier LLM. We present HALO (Hybrid Agent-Learned Orchestrator), which replaces it with a QLoRA-tuned 8B model (Llama-3-8B) trained on the agent choices of a prompted teacher, keeping only tra- jectories whose final plan an external validator accepts. This is outcome-filtered behavioural cloning: the validator certifies where a trajectory ended, not each step, yet it is enough to make imitation work, and removing the filter collapses performance. HALO places the policy behind three hardcoded rules for trivially decidable steps and selects among an expanded pool of 21 agents; the rules are essential, as the policy alone falls well short. Across 11 PDDL domains from PlanBench, Natural Plan, and classical planning, and measuring success as val- idator acceptance of the generated PDDL and plan, HALO matches the GPT-5- mini-orchestrated framework on PlanBench and classical planning, exceeds it by 37 points on Natural Plan, and stays within three points of Gemini-3-Flash, while cutting orchestration cost roughly 45× and 15× respectively and LLM calls per episode by 40–54%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.