acceptodds
Under review as a conference paper at ICLR 2027

Models as Evidence: When Collaboration Makes Bias Correctable

Abstract

Models reused to assess a population can inherit selection bias from the sources of their training data. We study how to recover the population rate, the positive-label rate before any source selects labels, when training data and selection logs are unavailable but each model's task and sources are recorded. Whether the population rate can be consistently recovered depends on how the models share sources. One additional model, trained only on existing sources, can make the rate consistently recoverable even when neither it nor the existing collection suffices alone. When several sources may be biased, we give an exact condition on these source records: with labels per source, the minimax mean squared error is when the condition holds and otherwise stays bounded away from zero, even when every model's fitted probability is observed exactly. A collection that fails this condition needs unbiased labels growing linearly with to reach that rate from a fixed batch of models. The condition extends to conditional predictions of models that approximate known source mixtures. When a reference task shares the selection, a criterion checkable before acquisition tells apart two candidates that match in source count, overlaps, and unbiased outputs: only one enables recovery. On human feedback under controlled selection, a candidate with worse predictions yields lower Brier loss once combined with existing models through source records, without calibration labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.