acceptodds
Under review as a conference paper at ICLR 2027

LatentCollab: Vision-Language Agent Collaboration Beyond Verbalization

Abstract

Vision-language agents extend vision-language models from isolated multimodal reasoning toward collaborative intelligence, but existing systems primarily exchange natural-language messages. This forces task-conditioned multimodal computation through a verbalization bottleneck that may omit fine-grained visual evidence, while KV-cache communication incurs substantial memory and communication overhead. We introduce LatentCollab, a latent collaboration framework that enables vision-language agents to exchange compact continuous messages without intermediate text generation. The sender recursively constructs a trajectory of final-layer hidden states. A lightweight translator, the only trainable component, maps the trajectory into a receiver-compatible continuous message, which is inserted into the receiver's input-embedding sequence after its visual tokens. Across seven benchmarks and three cross-scale or cross-family sender-receiver settings, LatentCollab outperforms receiver-only and text communication on all 21 pairs, and both KV-cache baselines on 18, with lower latency than all evaluated communication baselines. These results establish latent collaboration as an effective and efficient paradigm for vision-language agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.