Disagreement Has a Depth: Fail-Safe Consensus over Label Hierarchies
Abstract
Coordinated failures can flip a majority vote, yet sources that disagree on a label often agree on its ancestors: disagreement has a depth. We ask for the most specific statement that stays consistent with a strict majority of the uncorrupted sources, right or wrong, when up to of sources are corrupted. For label reports, the finest such rule, even among set-valued rules and rules that see the whole stream of examples, returns the deepest subtree holding more than reports. We derive its exact loss of depth, with a barrier at , and its exact risk under per-example and persistent adversaries, certifiable from i.i.d. labelled held-out data. For probability reports, backing off from the likeliest label to its first ancestor with mean probability at least is safe exactly when . Calibration to 95% coverage crosses this threshold on all four datasets, human labels included. Safe backing off is never finer than the finest safe subtree-valued rule for probability reports, so backing off matched to that rule's clean error crosses the threshold on all three machine-source datasets, with 1.2 to 2.7 times the rule's per-example worst-case error for two or three of ten corrupted sources. Safety costs abstentions and buys little against the non-adaptive failures we tested.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.