acceptodds
Under review as a conference paper at ICLR 2027

TRUE: A Trustworthy Reasoning and Unified Embodied Intelligence Benchmark for Multimodal Robotics

Abstract

Embodied vision-language-action (VLA) models are typically evaluated by whether their actions succeed, while the intermediate reasoning that connects perception, task progress, planning, and action remains largely unmeasured. We introduce **TRUE** (**T**rustworthy **R**easoning and **U**nified **E**mbodied Intelligence), a multimodal, multi-task dataset and benchmark for grounded and auditable reasoning in embodied intelligence. TRUE augments embodied trajectories with trustworthy reasoning-and-planning representations that describe task progress, intermediate sub-goals, scene-grounded evidence, and the consequences of recent actions while remaining constrained by the task specification and available observations. TRUE spans three tiers of increasing distance from standard robot learning: tabletop manipulation (CALVIN, LIBERO), CADRE-style lunar multi-rover navigation that we build in NVIDIA Isaac Sim, and real Perseverance Navcam/Hazcam imagery from NASA's Planetary Data System. The benchmark defines complementary two tracks under a fixed deployment interface: *control track* and *reasoning track*: the former measures whether reasoning supervision improves long-horizon action prediction, while the latter measures whether model-generated reasoning remains consistent with observable evidence across domains. We additionally provide a reference embodied model that combines visual observations, language instructions, explicit action history, and trustworthy reasoning within a shared representation. The benchmark reveals that (i) reasoning supervision and action history yield complementary gains in long-horizon control, improving CALVIN ABC→D average length by 38% over OpenVLA-OFT; (ii) gains vanish on near-saturated suites; and (iii) reasoning consistency degrades from simulated navigation (83%) to real planetary imagery (68%), exposing a substantial gap between simulated navigation and real planetary imagery. We release annotations, the Isaac Sim environments, evaluation code, and baselines. TRUE is designed to support joint study of embodied control, interpretability, and cross-task reasoning under a common benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.