Executable Intervention Games: From Attribution to Physical Consequences
Abstract
Explainable AI helps researchers formulate scientific hypotheses from predictive models, but fidelity to a model does not establish that the relationships it has learned hold in the physical system. Testing these relationships is difficult when input perturbations, such as masking, do not correspond to the physical actions of interest. To bridge this gap, we introduce executable intervention games, in which a model and a numerical reference answer the same questions about prescribed actions executed from shared starting conditions. Comparing their conditional action effects and Shapley attributions reveals response errors that inform researchers’ choices of action coverage or response supervision. In heat conduction, supervising single-action temperature differences improves held-out effect and attribution accuracy over longer field training and absolute-temperature supervision with identical labels. It also reduces average error in predicting how other active actions modify an effect in new materials, even though this dependence is not directly supervised. In cylinder flow, covering action combinations improves responses beyond controls with matched training states and histories, including at held-out Reynolds numbers. By preserving the physical question through explanation, feedback, and evaluation, executable games make learned relationships testable and support their improvement for scientific inquiry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.