RoboShackles: A Multilingual Safety Dataset for Human-Injury Prevention in Embodied Foundation Models
Abstract
Embodied Foundation Models (EFMs) integrate multimodal perception, future-state reasoning, and executable robot control. However, their safety alignment for preventing human injury remains largely underexplored, primarily because real-world data involving robots harming humans or creating hazardous household situations cannot be collected safely or ethically. To address this challenge, we propose a controllable safety-data construction pipeline, that transforms real-world DROID observations into safety-critical robotic rollouts. Specifically, we generate scenario images containing potential hazards, create hazard-triggering action prompts, and synthesize videos depicting hazardous events. Using this pipeline, a DROID-based robotic video dataset, RoboShackles, is constructed comprising 12,000 training clips and 1,200 hazardous multilingual evaluation clips across six safety categories. Based on this dataset, we conduct a comprehensive evaluation of fifteen representative EFMs and further align WoG with motion-stopping supervision for safety. Experimental results reveal that existing EFMs are highly prone to generating actions that could cause human harm. Moreover, fine-tuning with RoboShackles reduces the unsafe action rate from 98.92% to 9.67%, substantially improving the practical safety of EFMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.