CausGround: Causal Interaction Grounding for Adaptive Embodied Planning
Abstract
Multimodal large language models (MLLMs) have demonstrated promising reasoning capabilities, but deploying them for physical interaction requires reliable planning that accounts for when actions are executable and how they change the world. These preconditions and effects are often environment-dependent and are not fully captured by MLLMs, making adaptation to new environments challenging. Yet existing adaptation approaches either update model parameters through costly environment-specific training or rely on unstructured interaction experience that is difficult to ground and reuse. To address this challenge, we introduce CausGround, a non-parametric causal interaction grounding framework that distills interaction experience into reusable action-centered causal knowledge. CausGround learns from experience by identifying action preconditions and effects, and organizes them in a factorized world-state representation that jointly captures object attributes and inter-object relations. At planning time, task-relevant causal dependencies are grounded in the current observation, supporting backward reasoning over action preconditions and forward reasoning over action effects for action selection and multi-step causal planning. We demonstrate the effectiveness of CausGround across diverse long-horizon embodied settings spanning simulated high-level planning, interaction reasoning, and real-world robot manipulation. Evaluated on the adapted VirtualHome and RoboCasa benchmarks, CausGround consistently improves vision-based high-level planning across MLLM backbones, improving average task success from 24.9% to 53.3% on VirtualHome and closed-loop planning success from 43.0% to 57.0% on RoboCasa. Consistent gains are also observed on SWITCH benchmark, with average interaction reasoning accuracy improving from 25.8% to 31.9%, indicating stronger reasoning over action selection, recovery, and verification. Moreover, real-world experiments further show that CausGround can reason backward from MLLM-interpreted goal states to derive executable high-level action sequences for sequential manipulation. These results demonstrate that explicit action preconditions and effects provide a reusable causal grounding mechanism for long-horizon embodied planning across diverse environments without parameter updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.