Beyond Reasoning: Consensus-Guided Visual Collaboration in Latent Multi-Agent Systems
Abstract
Visual multi-agent systems (VMAS) leverage complementary reasoning roles for complex vision-centric tasks, but their prevalent reliance on natural-language communication creates an information bottleneck and substantial inference overhead. To address these limitations, we instantiate latent collaboration for VMAS, where agents reuse a shared visual KV cache and propagate intermediate reasoning KV states without generating textual messages. However, reasoning-state collaboration alone does not explicitly coordinate the role-conditioned visual attention patterns over the shared visual input. Our analysis shows that visual regions jointly prioritized across agents exhibit higher ground-truth region hit rates, suggesting cross-agent visual attention as a complementary signal for collaboration. Motivated by this observation, we propose V-Guide, a training-free framework that augments latent reasoning-state collaboration with explicit cross-agent visual coordination. Specifically, V-Guide first spatially marginalizes role-conditioned visual attention distributions to accommodate local token-level discrepancies, and then aggregates them into a geometric consensus distribution that captures cross-agent support. Finally, the resulting consensus distribution modulates the shared visual Value states, increasing the contribution of consensus-supported visual content for final reasoning. Extensive experiments across five visual reasoning benchmarks, four MLLM backbones, and different multi-agent architectures demonstrate that V-Guide outperforms text-mediated and pure-latent multi-agent baselines while retaining the efficiency advantages of latent collaboration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.