EmbodiedSafe: Benchmarking Context-Aware Runtime Hazard Handling in Embodied Agents
Abstract
Embodied agents driven by vision-language models (VLMs) integrate visual perception and language understanding to plan and execute multi-step tasks in interactive environments. However, deploying embodied agents in household tasks raises safety concerns, as unsafe behavior may cause property damage or physical injury. Existing benchmarks expose such behavior, but hazardous-execution rates alone do not capture appropriate responses across task contexts and stages. These responses depend on the current context, available corrective actions, and the timing of safety requirements. We present EmbodiedSafe, a benchmark for evaluating context-aware runtime hazard handling by embodied agents with and without runtime guardrails. EmbodiedSafe covers four evaluation levels: Hazard Rejection, Context Sensitivity, Hazard Resolution, and Hazard Prospection, assessing refusal, context-dependent decisions, corrective execution, and temporal safety. It contains 1,480 household evaluation samples covering 77 fine-grained hazard mechanisms and 20 temporal safety requirements. We evaluate five runtime guardrails across three embodied-agent workflows, with supplementary comparisons spanning four VLM backbones. Although guardrails reduce hazardous execution, failures persist across all four levels: the best context-pair accuracy remains below 58%, hazard-resolution success remains below 1%, and safe completion averages 0.22% under implicit temporal safety requirements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.