VIOLA: A Controlled Benchmark for Policy-Violation Diagnosis in a Multi-Agent LLM System
Abstract
Agent systems are increasingly deployed under explicit policies that say not only what they must accomplish but how. Task completion does not show that those policies were followed, and in a multi-agent pipeline a useful diagnosis must also name the requirement that was broken and the component responsible. Traces with such annotations do not occur naturally, and adding them by hand to runs of about 100k tokens is unreliable even for experts. We introduce VIOLA, a controlled benchmark for policy-violation diagnosis in the CUGA multi-agent system on AppWorld. A taxonomy of five categories and 12 violation types guides targeted interventions: one policy clause in one sub-agent's prompt is replaced with a contrary instruction, and two LLM-judge stages verify that the clause is gone and that the deviation appears in behavior. The result is 337 verified intervention traces over 11 violation types and five sub-agents, each labeled with its target type and agent, plus 63 controls. Judges from two other model families and seven human annotators characterize label reliability, and 78% of the intervention traces still receive a perfect task score, so outcome-based evaluation would miss them. We benchmark graph, sequence, and LLM detectors that see the agents' behavior but not their system prompts. A GCN attributes the violation to the right sub-agent in 48.9% of test traces and identifies its type in 26.3%, against 36.1% and 16.6% for the majority class. VIOLA supplies the annotated traces, the construction protocol, and the baselines needed to develop trace-based detectors, a first step toward automatic policy-compliance monitoring for multi-agent systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.