Safety-Sound Agent-to-Agent Handoffs under a Compromised Counterparty: Binding Attested Decision Bundles to Physical Safety Invariants
Abstract
Autonomous agents increasingly hand off safety-critical decisions to one another, and in cyber-physical settings those decisions drive actuators. Existing agent-to-agent (A2A) security certifies identity, provenance, and audit, while control-theoretic safety certifies a trusted controller; neither guarantees that the physical plant stays safe when the deciding agent is compromised. We introduce an invariant-binding risk contract that binds each admissible action to its certified worst-case actuator effect and admits it only if the recomputed successor state remains inside a certified-safe invariant set. Two results go beyond the prior runtime-assurance and secure-estimation lines: (i) a trust-boundary isolation—the identical runtime-assurance filter is fully breached over the sender-reported state yet safe over an independently attested state, so the trust boundary, not the filter, is decisive; and (ii) a composed guarantee preserving safety under a Byzantine sender and up to f of 2f+1 spoofed sensors, a regime neither line addresses alone. Underpinning both, a soundness theorem (Theorem 1) proves no accepted handoff can drive the plant into the unsafe set, for any behavior of an arbitrarily compromised sender, given a trusted verifier and a valid certificate—and the guarantee composes with learning: the certificate is learned (a monotone-by-construction network, continuum-verified before deployment) and safety holds even against a learned (cross-entropy-method) adversary or a real LLM from three providers (Groq, OpenAI, Anthropic). We authenticate the sender's decision bundle with HMAC while the verifier obtains its state from an independent attested sensor/estimation path—the signature secures the message, not the sender's state claim—and characterize the guarantee's boundary empirically: across cyber-physical triage scenarios the method attains a 0 unsafe rate under both a fully Byzantine and a learned adversary (claim-trusting and RTA-over-reported baselines fail 92–100%), at a tunable, genuinely-measured false-block cost, degrading gracefully as verifier integrity erodes. It generalizes to a non-monotone (band) barrier, the canonical adaptive-cruise-control CBF benchmark, and a 3-D coupled nonlinear regulator; on a non-monotone-disturbance plant a cheap corner reachability bound is breached (0.31) while our sound bound holds (0.000).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.