acceptodds
Under review as a conference paper at ICLR 2027

The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling

Abstract

Scalable analysis of long-form human–human, human–agent, and agent–agent interactions requires reliable supervision. Large language models (LLMs) provide a practical source of silver labels, but agreement among models can reflect shared labeling preferences rather than independent confirmation. We investigate this dependence in an enterprise pipeline that uses 76,000 silver-labeled customerservice interactions to train classifiers operating over tens of millions of conversations, with a separate human-labeled holdout of 924 examples. We introduce a gold-free audit: holding each labeler’s predictions fixed, we vary the model supplying the reference labels and measure changes in agreement without consulting human annotations. Across 48 model–prompt–rendering configurations, we observe source-dependent agreement that extends beyond exact self-comparisons to sibling models. In a representative comparison, agreement with sibling-model labels exceeds agreement with three cross-family sources by 3.7–3.9 percentage points. To mitigate this effect, we propose combinatorial silver-label construction, assigning each example to a randomly selected model–prompt tuple — randomizing the prompt as well, since prompt choice is itself a first-order driver of label quality. Under this construction, the source-specific agreement advantage seen with fixed-source labels is directionally reduced in our evaluation, though not to statistical significance at our sample size. The method distributes supervision across configurations while preserving exactly one inference call per example. Our central — and demonstrated — finding is that same-family consensus can reflect a source-dependent agreement pattern that resembles independent confirmation while not providing it. The gold-free audit makes this dependence measurable even where human reference labels are scarce, and randomized construction offers a practical, single-call route toward mitigating it in scalable conversation analytics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.