acceptodds
Under review as a conference paper at ICLR 2027

Beyond Single-Threshold Fairness: Residual Distributions and Group-Level Error Rates

Abstract

Fairness audits often compare groups at a single decision threshold, even though deployed thresholds can vary across settings and over time. We study how false-negative and false-positive differences between groups change over the full threshold range. Using the signed residual, we show that the Wasserstein distance between two groups' residual distributions equals the total area between their group FN and FP curves as the threshold varies. The two components show which error type drives the difference and how prevalence connects these group rates to FNR and FPR. We construct exactly calibrated populations that agree on the examined accuracy, ranking, calibration, average-error, and single-threshold fairness summaries while having different residual distributions and error gaps at other thresholds. Experiments on Adult, ACSIncome, HateXplain, and Measuring Hate Speech show that calibration can move group mean residuals closer to zero while residual distance increases. They also show that different prevalences can enlarge or reduce group FN/FP gaps. Across ACSIncome model configurations, validation FN and FP components preserve rankings of related test-time error differences more closely than fixed-threshold equalized odds or a mean-residual summary for the evaluated targets. Together, residual distributions provide a direct way to audit group error differences across decision thresholds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.