Truthful Data Procurement via Experimental Design
Abstract
Building a model for a specific task, whether reading medical images or training a warehouse robot from video, often requires buying data from multiple sellers. How should a buyer pay for data she has not yet seen? While the problem can be viewed as a reverse auction, the nature of data introduces three idiosyncratic difficulties. A buyer must evaluate data without fully accessing it, since once she has seen it she has no reason to pay (Arrow's inspection paradox). A seller who has already been paid can misreport labels. And even if sellers were honest, finding the welfare-maximizing set of datasets is computationally intractable. We study data procurement for supervised learning, where a buyer purchases datasets from strategic sellers with private costs, and we design a mechanism, the Greedy Auction for Data Elicitation via Surrogate (GRADES), which addresses those three difficulties. In the OLS V-optimal design setting of lu2024daved, GRADES is cost-truthful and individually rational in expectation over the label noise, makes truthful label delivery a dominant strategy, and achieves a bi-criteria welfare guarantee. We further extend GRADES to generalized linear models and to Bayesian decision problems with informational substitutes. Finally, we show empirically that it matches the welfare of an existing non-strategic procurement method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.