acceptodds
Under review as a conference paper at ICLR 2027

EmbodiedCausalEscape: Causal Reasoning under Embodied Interaction

Abstract

Causal reasoning in embodied environments requires agents not only to recover relations from observation, but sometimes to discover unknown action–effect mechanisms through intervention. In practice, however, the evidence needed for such discovery may itself be distributed across viewpoints or revealed only through interaction, making it difficult to separate failures of mechanism discovery from failures of embodied evidence acquisition. We introduce EmbodiedCausalEscape, a controlled 3D escape-room benchmark for studying causal mechanism discovery under varying embodied evidence conditions. It contains Association tasks, where task-relevant relations are encoded in observable evidence, and Intervention tasks, where an initially unknown action–effect mechanism must be identified through action. Matched diagnostic conditions preserve each underlying logical instance while varying mechanism information, spatial evidence access, and irrelevant interactions. Across four vision-language agents, Main-condition success ranges from 1/30 to 9/30, versus 25/30 for the human reference. On GPT-5.5, Intervention success rises from 4/15 to 14/15 when family-level mechanism structure is supplied, while spatial compaction raises overall success from 9/30 to 27/30. These results show that embodied task success depends not only on underlying task logic, but also on how mechanism-relevant evidence must be acquired and integrated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.