Translating Frontier AI Risk Narratives into Executable and Verifiable Testbeds
Abstract
As frontier AI systems shift from conversational assistants to autonomous agents with direct access to real environments, their safety risks have developed from toxic text generation to actionable system hazards. While recent studies reveal capable agents exhibiting dangerous strategic behaviors, such as peer preservation, covert code sabotage, and in-context scheming, these frontier risks remain confined to narrative reports without reproducible execution testbeds. We present VeriEnv, an end-to-end frontier threat modeling framework that synthesizes verifiable execution environments directly from natural language risk descriptions, enabling reactive red-teaming and causal attribution. VeriEnv isolates observable interfaces from backend ground truth, enforcing the Benign Solvability Axiom via closed-loop container validation to eliminate infrastructure false positives. Furthermore, VeriEnv formalizes dynamic red-teaming as Guarded Transition Systems under immutable cryptographic locks, casting safety oversight as discrete Bayesian causal inference over execution traces. Across ten documented frontier risk disclosures on six advanced models, VeriEnv achieves an 87.5% setup success rate, robustly elicits strategic vulnerabilities, and isolates actionable thresholds for defensive hardening.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.