TutelaBench: Evaluating the Safety of LLM Planners for Care Robots Among Simulated Humans
Abstract
Large language models are increasingly proposed as high-level planners for assistant robots, including in caregiving. Existing embodied safety benchmarks place the agent in a world without people, where safety is defined as “not doing dangerous things”. However, in a care setting, for instance, failure often looks different: something necessary is left undone, a commitment is forgotten, or the wrong person's demand prevails. To tackle this, we introduce TutelaBench, a benchmark that places an LLM-controlled agent in simulated elderly care and child care environments populated by scripted human characters who ask, insist, refuse, and need help. Requirements can come from multiple sources (deployer policies and the wishes of the persons cared for), and these sources can conflict. The world changes on its own clock, so that doing nothing can itself be unsafe, forcing the agent to revise and adapt the original plan. Each run yields a task score and an independent safety score. Safety is operationalized as compliance with explicitly stated requirements, which can be violated by action or by omission. These scores are computed from simulation state rather than judged by an LLM. On eight open-weight models and five agent architectures, TutelaBench shows task scores on the safety-critical scenarios barely separate models or architectures, whereas safety differs systematically. Stating the guideline in the prompt raises mean safety only moderately but increases the share of scenarios that are safe in every run by about half. A guideline critic raises safety further. However, it also uses about 2.5 times more LLM calls than a plain ReAct agent. Even the safest architecture violates a requirement in about a third of the scenarios over ten runs. The tested models and architectures thus often meet safety requirements, but none does so reliably across repeated runs, which leaves safe LLM-based planning for care robots an open problem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.