AI Assurance with Double Machine Learning
Abstract
AI assurance is the problem of testing whether artificial intelligence (AI) and machine learning (ML) systems adhere to a set of policies. Machine learning systems are particularly hard to assure since they are error-prone, unfair and biased. For example, is an auditor's claim that somebody is manipulating the system actually true or is there a more mundane explanation owing to an inaccurate or unfair classifier? To this end, we advocate the use of double machine learning (DML) as a framework for testing falsifiable assurance claims against ML systems. The approach directs ML upon itself to explain away its own workaday behaviors — behaviors that otherwise interfere with behavioral ML assurance tests. We study DML on the problem of detecting manipulation in content moderation systems. We find that DML allows us to isolate the effects of manipulation even in the presence of other interfering phenomena. On manipulation detection, DML achieves an AUC of 0.92 compared to fairness and statistical baselines which achieve AUCs around 0.76. We find DML succesfully detects a variety of manipulation attempts, including loss function manipulation, intentional bias amplification and the direct manipulation of its moderation decisions. Strikingly, in the latter case, DML estimates directly reflects the percent of manipulated content moderation decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.