Text-guided Active Image Selection for Radiology Report Generation
Abstract
Training vision-language models (VLMs) typically relies on large-scale paired datasets. However, na\"ive data scaling is often suboptimal or even impractical when one modality is significantly more expensive to obtain than the other. In medical imaging, for example, text reports are considerably cheaper to transfer and store compared to their corresponding imaging scans. In this work, we introduce the task of text-guided active image selection to efficiently construct a dataset for training report generation VLMs. Unlike standard active learning, where the pool consists of unlabeled inputs, we instead formulate the selection task in the text space. We instantiate this formulation with a suite of acquisition functions that identify the most informative reports for the downstream image-to-text task; their corresponding images are then queried to expand the training set. Experiments on two large-scale chest X-ray datasets show that text-guided selection achieves performance comparable to image-text selection, where both modalities are available. Moreover, we show that acquisition functions improve over random selection in terms of CheXbert-F1 at small data budgets. Finally, we demonstrate the feasibility of text-guided selection for cross-institution acquisition. These findings motivate rethinking data selection for medical VLMs by shifting the acquisition process from the expensive image space to the more accessible text space.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.