AEGIS: Evidence-Grounded Agent Safety Verification with Diffusion Scoring and Ising-Structured Risk Inference
Abstract
Reliable oversight of tool-using language agents requires contextual reasoning, auditable decisions, and low runtime overhead. We present AEGIS, a sidecar neuro-symbolic verifier that separates offline policy development from runtime detection. Rubric-specific extraction grounds twelve safety predicates in execution evidence. MaGiC uses frozen LLaDA-8B to score complete semantic verbalizers under bidirectional masked contexts and paired real/null contrasts. Ising-structured free-energy inference integrates these measurements into global risk; Policy As Code combines that score with hard constraints and abstention. We characterize when evidence contrasts and learned interactions improve their respective losses, and derive cascade score and threshold-decision bounds under explicit assumptions. On Agent-SafetyBench, Mixed-SafetyClaw, and Upward Deception, AEGIS achieves the highest precision on two benchmarks and competitive precision on the third among valid predictions, with – service-time speedups over direct and reasoning-augmented GPT-5.5 and Opus-4.8 judges. Component studies support moderate evidence correction and learned interactions, while recall gaps remain. The results establish a practical precision–latency trade-off with an 8B runtime model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.