Evaluating Safety of LLM Agents in Context
Abstract
AI safety benchmarks evaluate LLM-based agents with empty context, isolating the target behavior and thereby strengthening construct validity of the measurement. In real-world settings, however, autonomous and personal agents continually execute tasks over long horizons, accumulating input context from earlier interactions. Evaluating agent safety with prior context is costly and remains challenging because context can follow a variety of distributions and introduce dependencies between tasks. In this work, we study how the distribution and depth of prior context influence subsequent safety behavior across three benchmarks and nine frontier agentic models. We introduce a measure for meta-evaluating the sensitivity of agent safety benchmarks to contexts formed by executions of benign and harmful tasks. We probe its empirical range using behaviorally controlled contexts, such as executions compliant with unsafe requests, which are elicited through multi-shot jailbreaking and transferred across agents. Overall, we find that continual context shifts agent safety depending on its distribution and the effect often intensifies as context depth increases. Prior task executions frequently make agents overly cautious, resulting in overrefusal. However, safety deteriorates significantly when repeated compliance with unsafe instructions leads to emergent misalignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.