acceptodds
Under review as a conference paper at ICLR 2027

Iterative Deployment Improves Planning Skills in LLMs

Abstract

Frontier LLMs have become much better at end-to-end planning across recent model generations, but their training pipelines are not public, so the exact cause of this progress is unknown. Earlier attempts to improve planning skills of LLMs relied on planning-specific supervision, such as optimal plans from a classical planner, carefully shaped rewards, or structured validator feedback. We test whether a much weaker signal suffices: the model's own attempts, filtered by a pass/fail check of each final plan. Starting from Qwen3 4B Thinking or Gemma 4 E4B, we repeatedly prompt the model on PDDL tasks, check each final plan with a symbolic validator, keep the traces whose plans pass, and fine-tune the next generation on the pool of successful traces. Across three different domains, later generations solve substantially more held-out test tasks, including tasks from a harder parameter distribution, and find valid plans with longer horizons. Over ten generations, performance plateaus without model collapse. Our study also shows that curation is essential: training on all generated traces performs much worse despite using more data. Because we hold the base model, tasks, validator, and tuning recipe fixed, these results show that this minimal mechanism is sufficient to produce large planning gains. It is therefore a candidate explanation for part of the progress of frontier models, although our experiments cannot show how those models were trained.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.