acceptodds
Under review as a conference paper at ICLR 2027

A Minimal Agent for Planning Based on Reusable Domain Strategies

Abstract

In planning, an agent must compute a sequence of actions that transforms an initial state into a goal state. Frontier large language models (LLMs) now compete with state-of-the-art planners, but they are proprietary and expensive to run, while small open-weight LLMs remain behind. Planning problems, however, are organized into domains whose instances share the same actions and structure. In this paper, we exploit this structure with a minimal planning agent that adds two components to a small LLM, a domain strategy and feedback on its mistakes. The domain strategy is a natural-language text that explains how to compute a plan for any instance of the domain. The agent uses the strategy to guide the LLM, which returns a plan. A sound plan validator either accepts the plan or rejects it with feedback, and the LLM revises the plan with this feedback until the plan is valid or the call budget runs out. A frontier LLM writes the strategy once per domain, and the agent reuses it for every instance of that domain. We evaluate the agent with two small open-weight LLMs on the International Planning Competition (IPC) 2023 benchmark and on a benchmark of novel domains. The strategy increases the number of solved instances for both models, on both benchmarks, and at every budget. With up to five calls and feedback, GPT-OSS-20B improves from 28.3% to 40.6% on the IPC benchmark and from 0.4% to 17.4% on the novel benchmark. On the IPC benchmark, Gemma-4-31B solves 62.5% of the instances, more than a strong planner (56.7%) and GPT-5 (56.9%). To reach the coverage of GPT-5, it costs 9.8 times less than GPT-5 and 13.0 times less than Gemini 3.1 Pro. With feedback and up to 42 calls, Gemma-4-31B solves the same number of instances as the frontier model Gemini 3.1 Pro at a fraction of the cost. Overall, with only a strategy and feedback, small open-weight LLMs compete with planners and frontier models, and the computation spent once per domain is reused across all of its instances. This suggests that the same agent could be applied to problems beyond planning whose instances share a common structure from which a strategy can be defined.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.