acceptodds
Under review as a conference paper at ICLR 2027

Can Text Connect Images and Audio Without Paired Training Data? A Controlled Study

Abstract

Adding a new kind of data to an AI system usually requires paired examples and joint retraining. We test a simpler route. One stage learns from images and captions. A separate stage learns from audio and captions. We then compare images with sounds, even though these small connectors never trained on matched image–sound pairs. On a noisy 4,411-row proxy built from AudioCaps video thumbnails, the best system reaches a average retrieval score. That is about 34 times chance, but only of the score from ImageBind, a large system trained directly on image–audio pairs. The proxy contains a signal, but it is too weak for practical search. Changing the text model does not fix the gap. The data source and audio encoder can change results sharply, and a more flexible lightweight method does not win consistently. The lesson is simple: this staged approach can produce a weak match without matched image–audio pairs in its connector training, but text does not guarantee a useful connection. Builders must test their data, encoder, and evaluation pairs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.