Active Learning for Conditional Flow Matching Model in Engineering Design
Abstract
Although the flow matching model has demonstrated powerful capabilities in modern machine learning, its training notoriously relies on an incredibly large scale of high-quality labeled samples. Nevertheless, the acquisition of high-quality labeled datasets is hindered by exorbitant labeling costs in certain fields, notably engineering design. Therefore, selecting the most informative samples for training at minimal cost poses a key challenge in this field. This issue constitutes a central topic in active learning, a subfield of machine learning dedicated to maximizing model performance while minimizing annotation cost. The central challenge involves developing a query strategy to acquire the most informative data samples with minimal labeling effort. This paper presents a pilot study that investigates the application of active learning, which traditionally explored within the context of discriminative models, to flow matching models. By analyzing conditional flow matching through a piecewise-linear closed-form formulation, this work provides insight into how individual data points can affect the diversity and conditional accuracy of the resulting generative model. Building on this analysis, we propose two distinct query strategies: one aimed at enhancing model diversity, and the other designed to improve model accuracy. These strategies highlight a dataset-level trade-off between diversity and accuracy, and motivate a hybrid query strategy that balances the two objectives. Extensive experiments validate the effectiveness of our approach, showing that the proposed query strategies outperform baselines in terms of diversity and accuracy, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.