acceptodds
Under review as a conference paper at ICLR 2027

Language-Guided Safe Policy Optimization for Multi-Agent Locomotion control

Abstract

Can high-level language guidance improve safe coordination in multi-agent continuous control? We study this question in safety-constrained multi-agent reinforcement learning (MARL) for cooperative robotic locomotion, where multiple agents jointly control a shared body under an explicit safety budget. We introduce a framework that augments existing multi-agent policy optimisers with an episodic language-based meta-controller and reward penalisation. During training, the meta-controller periodically queries a large language model (LLM) using a compact textual summary of the system state and recent safety statistics to generate a cooperative strategy. The strategy is embedded and incorporated into each agent's observation through a learned gating module, while the underlying policy is trained under centralised training with decentralised execution (CTDE). Unlike Lagrangian constrained-MARL methods, our approach uses fixed reward penalisation and does not require additional cost critics or dual-variable updates. Experiments across six Safe Multi-Agent MuJoCo locomotion tasks show that language-guided policy optimisation improves the reward–safety trade-off across multiple robotic locomotion settings. In particular, LLM-HAPPO achieves higher rewards than standard constrained MARL baselines and a matched HAPPO variant using the same reward shaping, while generally maintaining low safety costs. The framework is also compatible with different MARL backbones. Overall, our results demonstrate the potential of episodic language guidance as a lightweight mechanism for improving safety-aware multi-agent policy optimisation without replacing the learned decentralised policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.