Stabilizing Multi-Agent System Evolution Through Evidence-Guided Revision
Abstract
Persistent multi-agent systems should turn experience across recurring tasks into reusable problem-solving procedures that improve over time. Yet every accepted revision is inherited by later tasks and future revisions, so harmful updates can compound alongside useful ones. The central challenge is determining what evidence justifies allowing a revision to shape the system's future. We introduce EMAS, which turns cross-task experience into evidence-gated MAS revisions while keeping the underlying LLM fixed. Its Step graph aligns diagnoses and edits; recurring diagnoses across tasks trigger bounded proposals, and paired validation on separate instances determines retention. After two epochs, validation-selected checkpoints improve task-weighted held-out test accuracy over the initial MAS across four benchmarks by 4.10 and 12.09 percentage points with Kimi K2.6 and Qwen3.6-27B, respectively. Across matched V0–V20 Game24 trajectories, Full EMAS yields the lowest cumulative regression for both backbones. On Qwen, allowing one task to trigger a proposal and removing paired Validation raise this burden to and the Full-EMAS level, respectively. These results show that evidence-controlled inheritance supports continued improvement while limiting damage along the evolution trajectory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.