ProcEvo: Self-Evolving Multi-Agent Orchestration via Process-Guided Attribution
Abstract
LLM-based multi-agent systems increasingly rely on learnable orchestration to coordinate specialized agents for complex tasks. Self-evolving orchestration further seeks to improve the orchestration policy from accumulated execution experience, making effective credit assignment over past trajectories critical. However, a single task- or trajectory-level outcome cannot distinguish the heterogeneous effects of individual agent transitions: the same execution may contain useful, redundant, and harmful transitions. Inspecting local state changes provides finer-grained information, but a state change alone does not reveal whether a transition constitutes meaningful progress, since repetition or unsupported reasoning may also alter the execution state. Therefore, we propose ProcEvo, a framework for self-evolving multi-agent orchestration that performs transition-level process attribution. ProcEvo first compares the task states before and after each agent invocation to estimate its local effect. It then uses an external LLM to extract evidence for a checklist instantiated with the task, applying relevance, consistency, support, and advancement to assess the evidence. The resulting process evidence is converted into transition-level process credit and combined with terminal supervision to optimize the learnable orchestrator. Across evaluated subsets of MMLU-Pro, GSM-Hard, and GPQA Main, ProcEvo achieves an average accuracy of 72.76%, compared with 67.47% for the highest-averaging competing baseline. Component ablations and process-level analyses further highlight the distinct roles of progress estimation by comparing task states and checklist verification. The implementation is available in the supplementary anonymous repository: https://anonymous.4open.science/r/ProcEvo-2EC7
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.