Human–AI Disparity Distortion: A Component-Referenced Diagnostic for Hybrid Decision Systems
Abstract
Fairness audits of hybrid human–AI systems often examine disparity in human judgments, algorithmic predictions, or final hybrid outcomes. Taken alone, these quantities do not distinguish the disparity implied by the component systems from non-additive disparity associated with their integration rule. We study this distinction through a component-referenced diagnostic. For an integration rule combining human probability and algorithmic probability , additive pooling yields group disparity exactly by linearity of expectation. We define the Human–AI Disparity Distortion Index (HDDI) as the difference between the disparity produced by and this additive reference. HDDI is identically zero under linear pooling; more generally, nonzero HDDI reflects group-differential departure from additive integration rather than nonlinearity alone. We evaluate HDDI across three complementary settings. On ACSIncome, controlled simulated human–AI pairings establish its behavior across models, groups, and integration weights. On Civil Comments, HDDI extends to identity-sensitive text moderation and is often more pronounced along a non-toxic, false-positive-relevant pathway; finite-sample recovery improves substantially with sample size, with median RMSE decreasing from at to at . On COMPAS, nonzero interior HDDI persists when human probabilities are derived from observed judgments rather than simulation, reaching , although its magnitude depends on how human judgments are aggregated. Across all three settings, HDDI complements rather than replaces conventional disparity and fairness measures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.