acceptodds
Under review as a conference paper at ICLR 2027

CALLOSUM: One Path between Seeing, Search, and Speaking

Abstract

Vision-language models (VLMs) increasingly serve as retrieval encoders in deployment systems, where their embedding space becomes a fixed interface for lookup collections, rerankers, and downstream applications. Yet naive generative fine-tuning can silently alter that interface: rank lists drift, thresholds shift, and candidates change despite stable aggregate retrieval scores. Can a unified VLM add generation while keeping retrieval drift within a measured budget? We propose Callosum, a one-path adaptation method that treats the original retrieval representation as an anchor. CAL–LO–SUM captures our method's three principles: preserving a cache-side anchor, adapting through low-rank operators, and unifying multimodal retrieval with generation. We show that Callosum shifts the retrieval–generation frontier beyond architecture-matched baselines and approaches the generation quality of a frozen-anchor-plus-generator reference. To examine system stability, we compare full rankings, top-10 candidate sets, and threshold decisions, finding Kendall , 93% top-10 overlap, and merely 3.2% threshold crossings. Thus, Callosum enables VLMs to see, search, and speak along one path while preserving the retrieval infrastructure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.