What Transfers from Synthetic Planning? A Controlled Study of LLM Post-Training
Abstract
Synthetic planning offers verifiable supervision for language-model agents, yet strong task performance leaves open what these agents learn and transfer. We introduce PlanWorlds, a controlled testbed that connects executable symbolic worlds with natural-language interaction, enabling systematic study of supervision interfaces, training trajectories, and learning objectives. Across two model families and six post-training strategies, we find that planning post-training produces distinct execution and transfer profiles. Holding expert solutions fixed, full-plan and interactive supervision lead to sharply different execution behaviors. Moreover, policies that achieve similarly strong, often near-saturated performance within PlanWorlds can diverge substantially on downstream planning tasks, and further post-training from the same SFT checkpoint can reshape transfer without changing training-environment success. We further show that these transfer gaps have different interpretations. Some largely contract under target-interface controls, while others persist and are accompanied by differences in robustness to local trajectory deviations despite nearly identical clean performance. These findings motivate evaluating planning transfer through complementary views of task success, interface compatibility, and sequential robustness, and enable controlled investigation of what language-model agents actually acquire from synthetic planning post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.