Safety Orchestration: Building a Routable Library of Atomic Safety Skills
Abstract
LLM agents need safety checks throughout the execution of complex, long horizon tasks. Open-source communities have developed many safety practices, but these are spread across separate artifacts that often provide the same functions. We screen 10,223 community artifacts and annotate 1,390 candidates to identify recurring safety functions. We consolidate these functions into a reusable library of 95 atomic capabilities across 19 functional archetypes. We introduce SAFETY ORCHESTRATOR, a training-free framework that selects phase-specific safety guidance and independently checks tool calls and outputs through host hooks. Across seven proprietary and open-weight models, it achieves average absolute reductions in Unsafe Action Rate of 15.42% on SABER and 8.84% on OpenAgentSafety, with an average absolute accuracy decrease of 1.28% on Terminal-Bench 2.1. Ablations show complementary benefits from guidance and host checks, with stage-specific reference selection outperforming loading the full library upfront in reducing unsafe actions. Evaluations across different agent runtimes also show that the same safety library can be reused. These results support modular runtime safeguards as a practical way to improve agent safety across models and runtimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.