Counterfactual Answer Prediction Training: Improving LLMs' Causal Understanding via Input-Perturbation Difference Prediction
Abstract
Large language models can solve math problems but may lack true causal understanding: models that correctly solve a problem often cannot predict how localized perturbations to it affect the answer. We introduce Counterfactual Answer Prediction Training (CAPT), a fine-tuning method that trains LLMs to predict answer deltas () under localized input perturbations, without re-solving the perturbed problem. A diagnostic study reveals a strong forward–counterfactual dissociation: Qwen3-8B solves 87.0% of GSM8K problems but predicts correct causal deltas on only 15.5% of relevant perturbations (dissociation score ), with the gap most severe for entity swap (10.0%) and condition negation (11.1%). We then construct a counterfactual training dataset of 4,817 triples from GSM8K and fine-tune with LoRA across four arms. Forward+CAPT (50/50 mix) substantially improves counterfactual competence (relevant CF accuracy: 15.5% 31.9%, , ), improves localization of causally irrelevant edits to 100% (from 97.8%), and improves out-of-distribution forward reasoning (MATH-500: 33.5% 42.0%, ). Critically, standard data augmentation fails to improve causal competence (18.1% vs. 31.9%, for the difference), indicating that the delta prediction objective—not mere exposure to perturbations—drives the gain. Scaling experiments across four model sizes (0.5B–8B) reveal a sharp capacity threshold: counterfactual competence remains flat at 9% for models up to 3B before jumping to 31.9% at 8B, while forward accuracy scales smoothly—providing additional evidence that forward and counterfactual reasoning are distinct capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.