Many Captioners or One? The Price and Payout of Mixing Frozen VLMs
Abstract
Recaptioning pipelines often mix several vision–language models (VLMs), on the premise that different models see different things. We test this premise against its obvious alternative, resampling one good captioner, at a matched number of captions. To this end, we build a controlled corpus in which 30 frozen VLMs caption the same M ImageNet training and k validation images under one prompt and one decoding specification, with resamples at identical settings. We find that mixing buys protection, not quality. In six of seven uses, from training a text encoder to auditing labels, the mix never reliably beats resampling the captioner that is best for that use, and it is worse in three, including text-encoder training, where much of the gap goes with caption length. Against a typical captioner, however, the mix gains in the same six uses: one captioner drawn at random per image trains a better text encoder than 29 of the 30 models used alone. This protection matters because the best captioner is hard to know in advance. It changes with the use and the prompt, its zero-shot accuracy says little about its value as a training source, and in the seventh use, a concept bottleneck, the captioner we declared best in advance is the weakest source we tested, so the mix wins. Mixing is therefore insurance against a poor choice rather than an improvement on a good one, at about twice the cost of generation. The code is available at https://anonymous.4open.science/r/ICRL_2027_34944_code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.