acceptodds
Under review as a conference paper at ICLR 2027

Concord: Co-Evolving Executors and Guards for Agentic Safety Rapport

Abstract

Large language model agents increasingly rely on external guards to supervise tool use and prevent unsafe actions. Yet guards and executors are often developed independently, creating a coordination gap when deployed together. We call the pair-level ability to enforce safety while preserving authorized task progress and recovering after intervention safety rapport. We identify three recurring failures: under-protection, where harmful behavior is allowed; over-refusal, where legitimate work is blocked; and coordination breakdown, where repeated Revise–retry interactions fail to resolve risk or advance the task. Evaluating existing guards with widely used executor models reveals all three failure patterns across deployed pairings. To specifically stress-test and quantify coordination breakdown, we introduce InterlockBench, which couples benign task requirements with malicious environmental injections and tests whether a pair can resist the injection while still completing the original authorized goal. We further propose Concord, a failure-attributed co-evolution framework that adapts the executor and guard as a deployed pair. Concord attributes unsuccessful interactions to the executor, the guard, or both, and evolves complementary procedural memories while keeping model weights fixed. Across multiple executors and agent-safety benchmarks, Concord improves task completion while maintaining high safety. On InterlockBench, it increases joint success while reducing both average Revise decisions and coordination breakdown. Joint adaptation outperforms one-sided alternatives, and the learned procedures transfer to new safety and general-purpose tasks. Memory analysis further shows that different executors induce distinct procedural guidance, highlighting the importance of adapting the deployed pair together.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.