LatentCollab: Vision-Language Agent Collaboration Beyond Verbalization
Abstract
Vision-language agents extend vision-language models from isolated multimodal reasoning toward collaborative intelligence, but existing systems primarily exchange natural-language messages. This forces task-conditioned multimodal computation through a verbalization bottleneck that may omit fine-grained visual evidence, while KV-cache communication incurs substantial memory and communication overhead. We introduce LatentCollab, a latent collaboration framework that enables vision-language agents to exchange compact continuous messages without intermediate text generation. The sender recursively constructs a trajectory of final-layer hidden states. A lightweight translator, the only trainable component, maps the trajectory into a receiver-compatible continuous message, which is inserted into the receiver's input-embedding sequence after its visual tokens. Across seven benchmarks and three cross-scale or cross-family sender-receiver settings, LatentCollab outperforms receiver-only and text communication on all 21 pairs, and both KV-cache baselines on 18, with lower latency than all evaluated communication baselines. These results establish latent collaboration as an effective and efficient paradigm for vision-language agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.