UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
Abstract
Large language model (LLM) multi-agent systems are usually assembled from prompts, tools, and hand-written control logic, while reinforcement learning (RL) is applied to one policy at a time. This mismatch makes it difficult to optimize intermediate roles, compare parameter-sharing choices, or reuse training infrastructure across workflows. We present UnityMAS-O, a general RL optimization framework that treats a user-defined multi-agent workflow as the unit of optimization. The framework exposes four first-class objects: logical roles, graph-structured trajectories, explicit role-to-model mappings, and role-, turn-, or trajectory-level rewards. A Ray-based controller executes arbitrary workflows and routes invocations to model-local worker groups built on vLLM rollout and PPO training. The same interface supports fully shared, partially shared, and fully separated parameters. We train and evaluate multiple workflows on retrieval-augmented QA and executable code tasks, demonstrating the effectiveness of UnityMAS-O across diverse multi-agent settings. We further explore multi-agent optimization under shared and independent parameter settings, examining their differences in task performance and training dynamics. UnityMAS-O provides a reusable framework for turning manually designed LLM multi-agent systems into trainable multi-agent policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.