Enhancing LLM Plan Diversity by Training on Multiple Partial-Order Linearizations
Abstract
Plans comprise steps that may be executed in more than one order to reach a goal, since pairs of steps may not necessarily be dependent on each other. LLM planning has become increasingly important, along with models' ability to emit valid sequences of steps and remain robustly invariant to their valid reorderings. However, such invariance requires producing similar likelihoods under different valid orders of steps, which is in general not true of pre-trained LLMs. In this paper, we train LLMs to increase the reliability and diversity of the plans they generate. During training, we feed them fixed or multiple orders of the same plan, either using standard language modeling or a plan-unshuffling objective. Likelihood comparisons and sampled generation on plans in multiple domains show that training on a single linearization of a plan's partial order reduces diversity among generated valid orders. By contrast, using multiple valid orders can vastly increase diversity among generated valid orders, which in turn could improve the ability of LLMs to avoid infeasible execution paths and parallelize independent steps in workflows. Finally, we find that the activations and behavior of the model reflect the underlying structure of the plans, with limited evidence that the likelihoods of step orders can be steered with causal interventions on the activations. Code: https://anonymous.4open.science/r/procedural_dags-1033
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.