EMBRACE: Embodied Models Benchmark for Robustness Across Cultural Environments
Abstract
Robots powered by foundation models are entering workplaces worldwide and are commonly evaluated on their ability to identify object types and ground them. However, recognising an object is not enough when the same visual cue can carry different meanings across jurisdictions. We introduce EMBRACE, a benchmark of ten safety-relevant object families across healthcare, manufacturing and retail, designed to test whether models adapt their interpretation to the stated jurisdiction. Across four vision-language models, the jurisdiction only partly changes behaviour as models tend to make the same selection for a given object family regardless of location. Shuffling jurisdiction labels within each object family preserves 64-82% of their original performance. When the locally correct object is unavailable, models often act on a convention used elsewhere rather than abstain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.