CausalScene: Explicit Intervention-Response Structure for Counterfactual Reasoning in Embodied Agents
Abstract
3D scene graphs capture where objects are, but not what happens when an agent acts on them. We introduce **CausalScene**, built on the *Causal 3D Scene Graph* (C3SG), which augments spatial scene graphs with directed, typed intervention-response edges (Aff, Fc, Tmp, Co), each carrying a strength estimated from simulation and validated through held-out intervention experiments. A **Causal Edge Predictor** (CEP) populates these edges by fusing cross-attention over the scene, a physics prior from 50,000 rigid-body simulations, and distilled LLM commonsense. A **Causal Query Language** (CQL) then delegates traversal to a graph engine, so every answer arrives with the intervention-response chain that produced it, making path accuracy measurable. On five 3D and physical-reasoning benchmarks (CLEVRER, CRAFT, ScanQA, SQA3D, EmbodiedScan) against 29 baselines, CausalScene leads on every one: 61.2% on CLEVRER counterfactual QA (+8.8 pts), 58.7% on CRAFT (+8.6), 54.3% on ScanQA (+5.1), 57.1% on SQA3D (+5.3), 62.4% on EmbodiedScan (+10.3), and 58.3% on RLBench (+5.6). On a real Franka Emika Panda across 120 tabletop trials, intervention consistency is 58.7% (vs. 63.4% on a held-out simulator), and human raters find the chains plausible (3.6/5) and correct (54.2%). Finally, we show that under 96% label skew the predictor collapses below a constant baseline and that a cost-sensitive two-stage detector recovers F1 to 0.687.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.