acceptodds
Under review as a conference paper at ICLR 2027

Causal Attribution from Video via Counterfactual Estimation

Abstract

Causal attribution from video asks whether an outcome would have changed if one of the observed actions had been different. Such attribution is essential for embodied agents that must reason about past actions to act safely next time. It is difficult because several concurrent actions can explain the same outcome. Pretrained models may be misled by correlations under confounding, and counterfactual outcomes are rarely labelled. We formulate video causal attribution as counterfactual inference given an observed video. To evaluate methods on this problem, we build TongSim-CF, a household video benchmark with controlled confounding from concurrent actions and scene conditions. We identify the resulting attribution target from repeated videos of the same environment without counterfactual supervision. We then estimate a local average of the target around the observed scene. Our estimator combines outcome predictions with a weighted residual correction derived from the efficient influence function. In the hardest simulated and real settings, it reduces the error of the strongest baseline by and .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.