acceptodds
Under review as a conference paper at ICLR 2027

Safety Orchestration: Building a Routable Library of Atomic Safety Skills

Abstract

LLM agents need safety checks throughout the execution of complex, long horizon tasks. Open-source communities have developed many safety practices, but these are spread across separate artifacts that often provide the same functions. We screen 10,223 community artifacts and annotate 1,390 candidates to identify recurring safety functions. We consolidate these functions into a reusable library of 95 atomic capabilities across 19 functional archetypes. We introduce SAFETY ORCHESTRATOR, a training-free framework that selects phase-specific safety guidance and independently checks tool calls and outputs through host hooks. Across seven proprietary and open-weight models, it achieves average absolute reductions in Unsafe Action Rate of 15.42% on SABER and 8.84% on OpenAgentSafety, with an average absolute accuracy decrease of 1.28% on Terminal-Bench 2.1. Ablations show complementary benefits from guidance and host checks, with stage-specific reference selection outperforming loading the full library upfront in reducing unsafe actions. Evaluations across different agent runtimes also show that the same safety library can be reused. These results support modular runtime safeguards as a practical way to improve agent safety across models and runtimes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.