acceptodds
Under review as a conference paper at ICLR 2027

CHARM: Benchmarking Capability-to-Harm Generalization in Generalist Robot Policies

Abstract

Generalist robot policies are increasingly capable of transferring learned skills across tasks and environments. Yet this generalization capability raises a critical and underexplored safety question: *Can robot policies trained on safe tasks generalize to harmful capabilities without task-specific adaptation?* We answer this question affirmatively by identifying a fundamental phenomenon: **capability-to-harm generalization**, where advances in robotic manipulation inadvertently enable generalization to harmful behaviors alongside benign ones. This finding reveals a critical decoupling between capability advancement and safety alignment in embodied AI systems. To systematically investigate this phenomenon, we introduce **CHARM** (**C**apability-to-**H**arm **A**ssessment of **R**obotic **M**anipulation), a benchmark designed to evaluate capability-to-harm generalization using matched benign–hazardous task pairs that rely on comparable manipulation primitives. CHARM operationalizes embodied semantic safety by translating high-level safety constitutions into structured hazard schemas. Leveraging an agent-driven generation pipeline in the high-fidelity OmniGibson simulator, CHARM automatically instantiates these schemas into diverse manipulation tasks and curates 50 representative scenarios spanning multiple hazard categories. We conduct comprehensive evaluations of 17 generalist robot policies, including modular policies, Vision Language Actions (VLAs), and World Action Models (WAMs), under a unified embodiment-aligned protocol without benchmark-specific policy adaptation. Our evaluation covers both explicitly harmful instructions and multimodal hazards, where risk must be inferred jointly from language and visual observation. Beyond terminal task success, we introduce the Hazard Progress Score (HPS), a state-based metric that quantifies progress toward hazardous outcomes with fine-grained precision. Our experiments reveal that 15 out of 17 robot policies demonstrate capability-to-harm generalization, exposing a critical safety issue in embodied AI. Safety alignment must advance in lockstep with capability development, and CHARM establishes the benchmark infrastructure essential to this goal. The project page is available at [CHARM-website](https://anonymous.4open.science/w/CHARM-website-61D6/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.