acceptodds
Under review as a conference paper at ICLR 2027

CabinSafeBench: Evaluating Context-Adaptive Safety in Intelligent Cockpit Agents

Abstract

Large language model (LLM)-based cockpit agents are increasingly capable of completing complex in-vehicle tasks, yet existing evaluations largely ask whether an agent can achieve a user goal, not whether it can adapt how that goal is achieved when safety-relevant conditions change. We study context-adaptive safety: the capability to preserve goal achievement, when a safe realization exists, while adapting execution behavior to the current driving context. We introduce CabinSafeBench, a benchmark based on controlled context intervention that keeps user intent and task goals fixed while varying safety-relevant world states. This design isolates failures of safety adaptation from failures of task understanding. We further develop CockpitWorld, a closed-loop environment coupling Android Automotive OS with CARLA, enabling safety to be evaluated from realized state transitions and execution consequences rather than tool calls alone. Across six LLM-based cockpit agents, task success is nearly saturated at 99.85%, while safe task success is only 75.38%. Under controlled context interventions, task success remains essentially unchanged, whereas safe task success drops from 99.81% to 40.19%. These results show that strong task competence does not imply context-adaptive safety: agents may know what goal to achieve while failing to adapt how it should be achieved as the world changes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.