Certified Execution for Agentic AI: Petri-Net Assurance Between Reasoning and Action
Abstract
Language-model agents ship code, call APIs and spend budgets, and a deployment cannot be recalled. We study an autonomous deployment agent that runs CI and a security scan, obtains human approval, then deploys, while an attacker hides an instruction in the issue description telling it to deploy at once. The difficulty is not one of accuracy. A guard that decides admission from the text it is shown cannot enforce “one approval authorizes one deployment”: every stateless monitor either admits an unauthorized action or admits none, enforcing the rule over a horizon needs at least monitor states, and over an unbounded horizon no finite-state monitor suffices. The guard must be able to count, and one place of a Petri net attains the bound exactly. We therefore hold the state in an agentic execution Petri net between the model and its tools, where approvals and budgets are marks that firing consumes, so one row of the state equation bounds deployments by approvals along every firing sequence. The synthesized supervisor is maximally permissive, as in classical supervisory control, and we prove that net-level safety transfers under four individually necessary assumptions. Empirically, four OpenAI models face seven indirect injections: unmediated, of injected episodes end in an unauthorized release. Under the kernel the models are hijacked just as often, no unauthorized deployment executes in any of kernel-mediated episodes, and pooled over the four models, task completion is higher than under any baseline, including no defense at all.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.