GenAL: Imbalance-aware Generative Active Learning for Discriminative Models
Abstract
Standard active learning (AL) frameworks for discriminative models struggle with imbalanced datasets due to misjudgment of sample informativeness under class imbalance. Current methods typically rely on algorithmic or sampling-based techniques to address imbalanced distributions, however, they overlook a more intrinsic solution: generative data augmentation. While traditional data augmentation approaches often fail in NLP tasks due to semantic loss, Large Language Models, with remarkable generalization and data synthesis capabilities, provide a promising pathway to mitigate class imbalance and further improve AL performance. In this paper, we introduce GenAL, a generative active learning framework robust across both standard and imbalanced datasets. To specifically address class imbalance, we design an imbalance-aware query strategy (BIRDS), coupled with a generative annotation framework (IGR) incorporating an automated refinement mechanism to synthesize diverse, high-quality data for AL training. Extensive experiments across both standard and imbalanced settings demonstrate that GenAL and BIRDS achieve superior performance compared to state-of-the-art query strategies, while IGR consistently delivers performance gains across different baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.