Zero-pool Active Learning: Choosing Experiments Through Metadata Alone
Abstract
When datasets are built through prospective generation, whether via physical experiments or complex simulations, the primary cost lies in producing the samples rather than annotating an existing pool. Active learning has the right objective for this setting but the wrong assumptions: it presumes an initial labeled set, a pool of unlabeled candidates to score, and a reliable post hoc oracle, none of which can be assumed when data must be generated. We formalize this zero-pool setting and introduce ZALMBIES, an algorithm that acquires data by specifying experimental conditions rather than by selecting instances. Because such experiments are defined by a small number of categorical choices, we place a hierarchical Bayesian ANOVA model over the metadata lattice, fit it to the primary model's realized out-of-fold failure, and interpolate between additive and multiplicative geometry with a Box-Cox power selected online. Candidate bins are scored by weighing predicted failure against an effective sample size that captures the statistical strength already borrowed from the surrounding lattice, and batches are built greedily with each selection depressing the scores of overlapping bins. We show that acquisitions are -optimal at cold start for the underlying linear model, and that under our exploration setting the resulting allocation asymptotically converges to Neyman proportions. Across four benchmarks spanning glucose simulation, quadruped robotics, wearable sensing, and sleep staging, ZALMBIES reaches the performance of random sampling using roughly half the budget. A sample here is an experiment, not an annotation, so the budget is spent in months of collection and the saving is one of schedule as much as of cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.