SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology
Abstract
Visual Multi-Agent Systems (VMAS) extend vision-language models (VLMs) through collaborative reasoning, but adapting agent perception and communication to the visual demands of each input remains challenging. Existing approaches often treat visual evidence acquisition, communication topology, and experience refinement separately, limiting how evolving skills can shape subsequent team behavior. We introduce **SkillGraph**, a framework for *vision-driven self-adaptation* that couples these processes through visual information. Given an image and a question, SkillGraph retrieves reusable skills to instantiate agents. A Multimodal Graph Transformer (MMGT) combines skill representations with question and image features to compute agent-specific selective attention and visually grounded representations. The resulting attention maps define agent-specific image crops, which are provided to the corresponding agents alongside the original image. The visually grounded representations further predict a directed communication graph, allowing visual context to guide both evidence acquisition and collaboration. During training, failed executions are diagnosed to refine a multimodal Skill Bank, whose updated skill representations are fed back into MMGT to reshape subsequent visual attention and communication topology. Experiments across four benchmarks, five representative multi-agent structures, and four VLM backbones show that SkillGraph consistently improves performance while reducing inference-time token consumption. Controlled analyses further validate the contributions of visual conditioning, agent-specific visual inputs, and skill evolution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.