Joint Supervision of Potential Outcomes for Causal Foundation Models
Abstract
Estimating treatment effects is central to science across many disciplines and requires counterfactual reasoning about alternative worlds. In the real world, however, each unit reveals its outcome only under the factual world, so the factual and counterfactual outcomes can never be jointly learned by any model trained on real data. The problem is therefore tractable only in synthetic settings, where both worlds can be generated. Recent progress on prior-data fitted networks (PFNs) has shown that models pre-trained purely on synthetic data can transfer to real-world prediction, achieving state-of-the-art tabular performance via in-context learning. This paradigm has since been carried over to causal effect estimation, predicting interventional outcomes given observational data. Despite this fully synthetic setup, existing causal PFNs carry the real-world constraint into their training objective, confining supervision to a single world's outcome per unit. We therefore propose a joint causal PFN training and inference architecture, apply it to existing causal PFNs, and show that it improves statistical efficiency and distributional information. Across semi-synthetic and synthetic benchmarks spanning diverse causal scenarios, jointly supervised models improve upon their baselines, establishing that joint learning is the preferable training setup for causal PFNs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.