Can Linked Binary Feedback Certify Ranking Disagreement?
Abstract
Several comparisons from the same evaluator reveal response patterns that pooled labels discard. Do these patterns establish different rankings, or can they reflect differences in decisiveness and dependence between answers? Under a logistic response model, every fixed finite linked-binary protocol on at least three stimuli admits a shared-ranking population and a population with a fixed positive fraction of ranking reversals that generate exactly the same joint responses. One distribution of decisiveness serves all comparison bundles. This rules out a uniform certification guarantee from finite linkage alone. We characterize when a fixed shared-ranking direction imposes testable restrictions on the joint response law. In a three-stimulus benchmark, near ties reveal a signal quadratic in ranking heterogeneity, but dependence of the same order can exactly reproduce the complete repeated-comparison law under a shared ranking that permits ties. For a calibrated benchmark with a known dependence budget, matching detection bounds separate the effects of sampling noise and residual dependence. Computations corroborate these transitions. A controlled study of language-model judges illustrates the diagnostic: all three tested models, each using a fixed six-persona roster, violate necessary shared-order conditions on a preselected narrative-rhythm triple at simultaneous 95% confidence. The results connect informative comparison designs to the dependence calibration needed for ranking-disagreement certificates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.