Staged Latent Chain-of-Thought for Sequential Counterfactual Prediction with Universal Holdout Calibration
Abstract
In personalized marketing and medicine, interventions shape both immediate outcomes and the states that influence future responses. Evaluating dynamic policies therefore requires predicting how state and outcome trajectories evolve under alternative intervention strategies. Long-term randomized holdouts are already used in practice, particularly in personalized marketing. Simply pooling these observations with intervention histories increases sample size without explicitly correcting hidden-confounding bias in counterfactual predictions. We develop Staged Latent Chain-of-Thought with Universal Holdout Calibration for this sequential counterfactual prediction problem. Our approach combines randomized evidence from long-term holdouts with explicit supervision of intermediate latent computation. First, we use randomized comparisons within subgroups defined by pretreatment covariates to construct causal anchors for state and outcome trajectories, providing policy-effect calibration targets without requiring all confounders to be observed. Second, we structure latent chain-of-thought around the progression from states to outcomes to effects, supervising successive latent stages with state and outcome prediction targets and calibrating effects derived from the final counterfactual rollouts. Experiments on a CVSim-based benchmark demonstrate that our method outperforms state-of-the-art baselines in effect estimation and policy-outcome prediction. Holdout calibration alone reduces effect RMSE by 30.1% and policy-outcome RMSE by 21.6% relative to a G-Transformer trained on pooled treatment and holdout data. Staged supervision further reduces effect and policy-outcome RMSE by 2.58% and 1.94%, respectively, relative to the holdout-calibrated G-Transformer, while outperforming looped Transformers and unstaged latent chain-of-thought. This work introduces a new framework for sequential counterfactual prediction under dynamic treatment policies, transforming randomized long-term holdouts into trajectory-level causal supervision and structuring latent computation around the progression from state prediction to outcome prediction and effect estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.