acceptodds
Under review as a conference paper at ICLR 2027

Auditing Accuracy Change in Human–AI Decision-Making: Partial Identification and Behavioral Constraints

Abstract

Research on human–AI decision-making aims to help people make more accurate decisions, but how much improvement after AI advice can existing studies establish? The data reported in these studies may include AI recommendations and final human decisions but may not determine this change when people's independent judgments are missing. We thus provide an analytical lens for evaluating how information and assumptions about human behavior constrain the possible accuracy changes. When people either retain their initial judgment or adopt AI advice, these records leave a range of possible accuracy changes whose width equals the final human–AI agreement rate and cannot be narrowed by collecting more data of the same kind. We examine two ways to narrow these bounds: sparse measurement of initial judgments and behavioral restrictions based on patterns or models. Evaluations on six human-AI decision tasks show that, behavioral patterns narrow the bounds, while the additional contribution of behavioral models varies across tasks. The relative precision of direct measurement and behavioral constraints depends on the task and measurement budget. Using this lens, we audit 94 claim–workflow instances from 56 studies and find sufficient evidence for the conclusions in 80 instances; 14 require additional information or assumptions for stronger interpretations. We provide implications for evaluation and interpretation in human–AI decision-making.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.