acceptodds
Under review as a conference paper at ICLR 2027

LEARNING TO ORCHESTRATE WITH COUNTERFACTUAL PREFERENCES

Abstract

Large language model (LLM) orchestration requires jointly deciding which workers to call, what subtasks to assign, and how to pass information between them. Terminal task rewards make this hard to learn, since one outcome must be credited across many coupled choices. We introduce CPO-Orch, an offline preference-learning approach that localizes this supervision. It constructs decision-localized counterfactuals by editing exactly one decision in an executed workflow (a worker, a subtask instruction, or an access edge), re-executing both variants, and retaining pairs whose outcomes differ consistently under repeated execution. It combines these with natural preference pairs, structurally matched to suppress worker-identity shortcuts, and trains an 8B orchestrator with direct preference optimization, without online rollouts. Across four reasoning domains, CPO-Orch reaches 85.56% macro accuracy, outperforming the strongest fixed worker in its pool (83.18%), the strongest single model evaluated outside the pool (82.97%), the best fixed multi-agent scaffold (84.02%), and an oracle that selects the best worker per domain, while ablations show that natural and counterfactual pairs are complementary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.