acceptodds
Under review as a conference paper at ICLR 2027

Large Language Models as Policy Optimizers for Cooperative Multi-Drone Payload Transport

Abstract

Cooperative transport of a cable-suspended payload requires multiple drones to coordinate formation, obstacle avoidance, and stabilization under coupled dynamics. Learning and adapting these behaviors is challenging when control decisions cannot be readily inspected or revised. We develop a framework that uses pretrained large language models (LLMs) to optimize a shared control program from feedback on simulated transport tasks. The optimizer alternates code revision with local parameter search and compares candidates on the same task instances. The simulator and low-level flight controller remain fixed, and deployment executes the frozen program without online language-model inference. We construct a benchmark covering standard transport, commanded emergency setdown, and stress conditions. Our method achieves higher overall task success than proximal policy optimization (PPO) on this benchmark under the reported training budgets. Optimization-flow ablations support combining diagnostic execution feedback with local parameter search on the main benchmark.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.