One Goal, Many Commands: Characterizing Denylist Fragility in AI Agents
Abstract
The adoption of AI agents is increasing rapidly. Terminal AI agents, i.e., AI agents that run in terminal environments, are widely used. Terminal AI agents rely heavily on shell commands to interact with host systems. They adopt denylists to mitigate the security risks introduced by executing arbitrary commands, which serve as the last line of defense against dangerous and harmful operations. However, denylists can be fragile, as the same goal (e.g., performing a harmful operation) can often be achieved by many commands. This paper presents the first systematic characterization of command denylist fragility in terminal AI agents. The paper formalizes the command denylist fragility problem and proposes an LLM-driven pipeline, ShellSieve, to detect such fragility. It prompts the LLM to propose possible bypasses and iteratively repairs them using feedback from a validator that executes them in a sandbox. In the evaluation, we applied ShellSieve to 4,668 real-world command denylists (containing 36,846 denylist rules) collected from GitHub. The evaluation shows several key findings, including that 69.5-99.8% of the denylists are fragile, that this fragility occurs consistently across projects and agents, and that one cause of this fragility stems from denylist authors overlooking less-known commands. Our pipeline and findings will hopefully facilitate future research and practice regarding the command denylists used by AI agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.