SafeHousehold: Benchmarking Vision-Language-Action Models for Household Manipulation with Implicit Risks
Abstract
Household robots must perform everyday tasks safely in visually rich environments, yet task instructions often omit risk conditions that determine the appropriate action sequence. We define an implicit risk as an instruction-omitted, visually observable condition whose presence requires a corresponding mitigation. For example, given the instruction to put an apple on the plate, a policy should clean a dirty plate first but place the apple directly when it is clean. Existing benchmarks do not jointly evaluate whether closed-loop VLA policies infer such risks, execute mitigations in the required temporal order, and avoid unnecessary responses when risks are absent. We introduce SafeHousehold, a household manipulation benchmark in which policies receive only the task instruction, RGB observations, and a proprioceptive state, without risk labels, mitigation descriptions, or safety specifications. SafeHousehold contains 35 task templates across seven risk categories and 1,750 manually reviewed human-teleoperated demonstrations. We evaluate four representative VLA models under single-risk, two-risk composition, and no-risk scenarios. The protocol distinguishes task success from safe task success with process-level metrics, using automatically computed scores for reproducible model comparison and human-reviewed diagnostics for failure analysis. In the single-risk scenario, policies often initiate a risk response without completing the mitigation; even the best-performing model achieves 39.0% task success but only 27.3% safe task success across seven risk categories. Safe task success is nearly absent in the two-risk composition scenario, while unnecessary risk responses remain frequent in the no-risk scenario. These results identify mitigation execution, responses to multiple risks, and risk-conditioned behavior as key bottlenecks in current VLA policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.