Defining and Benchmarking Active Data Labeling
Abstract
The availability of LLMs and VLMs enables the study of a new kind of machine learning setup, which we call Active Data Labeling (ADL): given a fixed set of data points, a novel classification task, and an oracle providing labels to this task, the goal is to annotate the entire set with oracle quality, but using as few calls to the oracle as possible. ADL is related to Active Learning, but also substantially differs from it: changing the goal from creating a model to annotating a fixed data set gives rise to a new set of affordances, methods and tradeoffs that could be explored (e.g., deferring certain classes of hard samples to the oracle while focusing on the easier, manageable ones). The ADL setup arises naturally with LLMs: a user asks to classify a set of documents or records by applying an LLM prompt to each one, and our goal is to provide a high quality annotation while minimizing actual LLM calls. We formalize the ADL learning setup, propose standardized evaluation metrics, and provide TeLa-Bench, a benchmark of cases on which it could be studied. We also present baseline algorithms that establish a lower bound on what can be achieved in the black-box ADL setup: a reduction of 43% in LLM calls while maintaining 98% agreement with the oracle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.