Human Before the Loop: Constraining the Action Space Improves Security, Reliability, and Capability in AI Agents
Abstract
Multipurpose AI agents are increasingly common, and so are the data breaches they cause. These agents are vulnerable to an attack called prompt injection, which can instruct a large language model (LLM) to find and deliver sensitive data to an attacker. Current proposed defenses carry their own trade-offs. Post-training a model to refuse malicious instructions fails to predict novel or unseen attack vectors. Live supervision slows down runtime operation. Security middlewares can only stop attacks they are built to recognize. Defenses that restructure an agent’s control flow have been documented to have drops in benign task completion. To address this, we present Human Before the Loop, an agent framework that determines the action space and data access in the design phase, leaving no runtime path open to sensitive data or destructive actions. Rather than relying on traditional open-ended tool calling, our framework uses a method we term semantic action abstraction, which navigates actions within a Hierarchical Extended Finite State Machine (HEFSM). This eliminates the chance of data exfiltration while increasing benign task success and realized capability. In our evaluations, our framework prevented all prompt injection attacks, whereas the baseline tool-calling agent failed to prevent 8.5% of attacks and delivered sensitive data to the attacker. Benign task completion rose from 76.5% with the tool-calling agent to 81.4% with our framework. Additionally, our framework enabled a 30B parameter model to complete the Pokémon Red tutorial phase in 341 turns, while the same model using a tool-calling harness made no progress beyond the initial stage across 1,000 model calls. That 30B model finished the tutorial in 22% fewer model calls than the larger model it was distilled from. This method supports the creation of autonomous agents that can be trusted to run fully unsupervised, without a trade-off between security and capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.