OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
Abstract
AI agents are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, tools, and execution backends act over shared user or enterprise environments. In such systems, safety becomes a system-level action-governance problem: whether a concrete pending action should be committed given policy-relevant state accumulated across a session. Existing safeguards operate at useful but fragmented boundaries, such as prompts, tool calls, GUI actions, or individual agent runtimes, making it difficult to enforce shared policies over composed action flows across heterogeneous execution paths. We present OpenAgentFlow, a control-plane/action-plane architecture that establishes the action-commit boundary as a shared enforcement interface. GUI, API, tool, and LLM-generated actions are normalized into a common AgentEvent stream and mediated by a shared pre-execution Policy Enforcement Point (PEP), while provenance, session state, audit evidence, and updatable policies are maintained outside individual agents. We evaluate OpenAgentFlow through a series of complementary system evaluations spanning controlled action-flow tests, a public external benchmark, post-deployment policy updates, and real Android execution. On a 300-case controlled suite, OpenAgentFlow achieves 94.00% accuracy and a 95.35% attack-block rate. On the complete 1,220-case AgentDojo-Traj split of TS-Bench, it achieves 97.62% accuracy, 96.59% unsafe-action recall, and a 1.96% safe false-intervention rate. Newly installed control-plane rules take effect without modifying protected agents, and the same enforcement path operates across live GUI, API/tool, and LLM-planned Android execution. These results show that a shared action-commit boundary provides a practical basis for consistent system-wide governance across heterogeneous agent execution paths. Code is available at https://anonymous.4open.science/r/Openagentflow-B161.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.