CoRefine: Joint Reconfiguration of Concurrent Agent Workflows
Abstract
Multi-agent systems powered by large language models (LLMs) have demonstrated strong capabilities in solving complex tasks through coordinated reasoning and execution. However, existing approaches largely optimize individual workflows or coordinate multiple workflows through scheduling and model routing within prescribed structures, leaving runtime joint reconfiguration under shared resource constraints underexplored. We propose CoRefine, a runtime orchestration system that jointly optimizes the structure of concurrent agentic workflows to balance quality, latency, and token cost under shared resource constraints. CoRefine first narrows the search space by evaluating each workflow’s candidate configurations according to their expected quality and execution cost. It then jointly evaluates the remaining combinations, accounting for the additional waiting time that workflows impose on one another through shared inference resources, and selects the combination with the highest estimated overall utility. In the evaluation run across 1,024 BigCodeBench episode contexts, CoRefine achieves the highest average utility among the evaluated methods. Compared with the fixed-workflow baseline, CoRefine reduces token usage by 68.8% and flow time by 35.3%, while achieving an accuracy of 81.0% versus 80.0%. These findings highlight workflow structure as an important control for allocating inference computation efficiently under shared resource constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.