Concord: Co-Evolving Executors and Guards for Agentic Safety Rapport
Abstract
Large language model agents increasingly rely on external guards to supervise tool use and prevent unsafe actions. Yet guards and executors are often developed independently, creating a coordination gap when deployed together. We call the pair-level ability to enforce safety while preserving authorized task progress and recovering after intervention safety rapport. We identify three recurring failures: under-protection, where harmful behavior is allowed; over-refusal, where legitimate work is blocked; and coordination breakdown, where repeated Revise–retry interactions fail to resolve risk or advance the task. Evaluating existing guards with widely used executor models reveals all three failure patterns across deployed pairings. To specifically stress-test and quantify coordination breakdown, we introduce InterlockBench, which couples benign task requirements with malicious environmental injections and tests whether a pair can resist the injection while still completing the original authorized goal. We further propose Concord, a failure-attributed co-evolution framework that adapts the executor and guard as a deployed pair. Concord attributes unsuccessful interactions to the executor, the guard, or both, and evolves complementary procedural memories while keeping model weights fixed. Across multiple executors and agent-safety benchmarks, Concord improves task completion while maintaining high safety. On InterlockBench, it increases joint success while reducing both average Revise decisions and coordination breakdown. Joint adaptation outperforms one-sided alternatives, and the learned procedures transfer to new safety and general-purpose tasks. Memory analysis further shows that different executors induce distinct procedural guidance, highlighting the importance of adapting the deployed pair together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.