acceptodds
Under review as a conference paper at ICLR 2027

PhysCogSafe: Diagnosing Physical Safety Cognition in Generalist Robot Policies

Abstract

Generalist robot policies, particularly Vision-Language-Action (VLA) models, are emerging as a central paradigm for robotic manipulation, making their safe operation in the physical world a critical concern. Existing benchmarks primarily characterize whether executions are safe, but provide limited evidence about whether observed safety reflects risk-sensitive physical cognition or incidental success. To address this gap, we introduce PhysCogSafe, a controlled diagnostic framework for safety-relevant physical cognition in robotic manipulation. Rather than treating safety as a binary endpoint, PhysCogSafe diagnoses whether a task-capable policy selectively adapts its behavior to physical risk along three dimensions: geometry-conditioned reasoning, property-conditioned reasoning, and multi-subgoal compositional reasoning. Each diagnostic probe holds the instruction and task objective fixed while forming a matched family comprising a benign baseline, a physical-risk condition, a control that retains comparable scene changes without the causal hazard, and an independently verified safe reference. Counterfactual contrasts and trajectory-level monitors jointly assess task completion and targeted harm. We instantiate PhysCogSafe in LIBERO, evaluating four representative VLA policies and one world-action model across 11 scenario families. Across 50 of 54 policy–scenario pairs, introducing a physical hazard increases targeted harm relative to an otherwise comparable scene in which the hazard is absent, supporting the intended risk-specific comparison. After controlling for underlying task capability, we observe that successful, harm-free completion falls from 69.5% without the hazard to 11.7% with it, while 66.2% of hazardous-scene outcomes involve harm. Failure modes also depend on the required physical reasoning: policies often complete tasks requiring property-conditioned reasoning unsafely, whereas tasks requiring multi-subgoal compositional reasoning more often end in both task failure and harm. Overall, none of the five evaluated policies reliably translates task competence into risk-aware physical behavior. These findings highlight safety-relevant physical cognition as a capability that should be explicitly addressed in the development and evaluation of generalist robot policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.