ViLaMAS: Latent State Handoffs for Multi-Agent Long-Video Understanding
Abstract
Multi-agent systems (MAS) for long-video question answering face a key bottleneck: progressive accuracy degradation across successive agent handoffs. Textual handoffs can omit visual evidence and intermediate reasoning, which may cause additional collaboration stages to undermine earlier understanding and limit the effective use of more elaborate MAS. We introduce ViLaMAS, a training-free framework for reducing this degradation within existing MAS workflows. Agents sharing a compatible frozen backbone transfer their complete accumulated multimodal KV caches and continuation metadata, allowing subsequent agents to continue autoregressive latent reasoning. Evidence-guided revisiting augments the inherited state with newly sampled frames and selected original frames, followed by latent updates that integrate the additional evidence. We also introduce Self-Verification Inference (SVI) as a baseline for long-video question answering. With Qwen3-VL-32B, SVI exceeds the reported accuracy of GPT-4o on EgoSchema and LongVideoBench. SVI with two textual agent handoffs reduces accuracy from 66.0% to 47.5% with Qwen3-VL-8B on LongVideoBench. ViLaMAS raises accuracy to 61.8% within the same collaboration structure. Across six MAS architectures, ViLaMAS reduces this gap by more than 60% on average compared with textual handoffs. These results support preserving and extending multimodal context to retain video understanding accuracy across compatible MAS.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.