Learning Probabilities of Causation with Mask-Augmented Data
Abstract
Probabilities of causation play a central role in modern decision making. Tian and Pearl first introduced formal definitions and derived tight bounds for three binary probabilities of causation (PoCs), such as the probability of necessity and sufficiency (PNS). However, estimating these probabilities requires both experimental and observational distributions specific to each subpopulation, which are often unreliable or impractical to obtain from limited population-level data. To solve this problem, we propose a machine learning (ML) framework that learns from a small set of reliable subpopulations and predicts PNS bounds for all subpopulations. We further introduce a mask-based augmentation strategy that augments the data and enables prediction of masked subpopulation queries. In experiments across five ML models, four Structural Causal Models (SCMs), and four data budgets, mask-augmented neural models consistently achieve the lowest average errors. Mask-Transformer and Mask-MLP obtain average MAEs of 0.0290 and 0.0341 on masked queries, compared with 0.0985 for the plug-in baseline, and also improve exact-query MAE from 0.1866 to 0.0469 and 0.0538. These results demonstrate both the feasibility of ML models for learning PoCs and the effectiveness of mask-based augmentation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.