Behavior Turing Test: Evaluating Human-Likeness of Humanoid Behaviors in Context
Abstract
Recent advances in humanoid robotics have enabled increasingly complex behaviors in real and simulated environments. However, existing evaluation remains largely motion-centric, focusing on task success, tracking accuracy, or motion-level human-likeness, while leaving a key question unexplored: would a human behave this way under the same task and scene? To address this, we introduce the Behavior Turing Test (BTT), which evaluates human-likeness at behavior level under explicit task and scene context. We construct HuCo, a dataset of 1,280 behavior clips spanning real humans, digital humans, real humanoids, and simulated humanoids, each paired with a Task Goal and Scene Description. We introduce Context-SMPL to suppress embodiment-specific appearance cues while preserving motion and scene context, and collect human judgments of contextual human-likeness. Our analysis shows that some humanoid behaviors can pass the Motion Turing Test yet fail BTT once task and scene context are considered, and that a clear human-humanoid gap persists under contextual evaluation. Building on HuCo, we establish HuCoBench for BTT and construct task and scene counterfactual subsets to evaluate context sensitivity. We further introduce CoBE, a Context-conditioned Behavior Evaluator that retrieves context-relevant behavioral evidence and achieves the strongest alignment with human judgments among evaluated methods. We will release dataset and code to support the community.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.