Towards Multimodal Generative Active Learning: Efficient Learning via Uncertainty Dynamics
Abstract
Active learning (AL) aims to improve data efficiency by selectively querying the most informative samples for annotation in deep learning. However, existing AL research focuses predominantly on unimodal and multimodal discriminative learning, leaving multimodal generative learning largely unexplored. We introduce the first framework for , where the learner must actively acquire textual ground-truth for multimodal inputs rather than conventional labels. This setting differs fundamentally from existing multimodal AL paradigms, which typically rely on predefined labels obtained through annotation, an assumption that does not hold in multimodal generation. To this end, we propose an innovative AL algorithm that jointly exploits uncertainty and diversity through uncertainty dynamics, facilitating effective and representative sample selection with improved acquisition efficiency. Empirical results on benchmark datasets demonstrate that our approach substantially improves annotation efficiency, achieving comparable performance with up to 33% fewer annotated samples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.