The Safety Criticality Number: Cascade Stability and Intervention Thresholds for AI Agents
Abstract
Risky actions by an AI agent may arrive in cascades: one event changes context or shared state so later events become more likely. We model this with a marked Hawkes process and separate two conflated quantities—an intrinsic cascade criticality (whether an uncontrolled cluster dies out) and operational monitorability (observation, containment, background risk, capacity). With observation probability and conditional containment , we prove within the model that control is subcritical exactly when ; we extend this to interacting agents via a controlled branching matrix and establish a tight testing rate in labeled parent events. Empirically, in a synthetic decision setting with all tools disabled and labels never executed, a prespecified auditable paired replication raises the downstream terminal unsafe rate from to over matched restrictive cases when the erroneous upstream conclusion is supplied (pp; McNemar exact ; all prompts and raw responses retained), while a multi-model panel shows heterogeneous per-edge transmission ( from 0 to ) and a realized cascade shows finite-horizon growth versus persistence consistent with . A second pre-registered study on a different task family observes no transmission on two current models; because the studies differ in more than verifiability, this contrast does not isolate verifiability as the cause and holds for the tested tasks only. A separate read-only trace study finds excess adjacency of elevated-access tool calls after conditioning on task and phase (Holm )—temporal dependence, not evidence that deployed-agent harms follow a Hawkes process. Finally, a prospective test—predictions fixed before each corrected collection on a previously examined case pool—forecast that placing an independent reviewer *last* is riskiest; the order held in both error-making models, resolved within one and pooled (McNemar ) but not the other, with a permutation-invariant baseline tying all placements. These are model-specific translations of a classical stability boundary, not a measured branching ratio or a deployment guarantee.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.