AgentRecover: Coordinated Interception and Recovery for Agent Safety and Utility
Abstract
Trustworthy large language model (LLM) agents must avoid unsafe actions (*Safety*) while completing legitimate user tasks (*Utility*). Runtime guardrails improve safety by intercepting unsafe calls, but interception alone can leave legitimate tasks unfinished and reduce utility. We present *AgentRecover*, a runtime method that coordinates interception and recovery to improve agent utility without relaxing safety requirements. Its central idea is to equip guardrails with precise feedback that supports Certified Recovery. AgentRecover introduces (1) *Guardrail with Feedback*, which intercepts unsafe calls and provides precise feedback, including root causes and permitted data sources, to support recovery, and (2) *Certified Recovery*, which constructs a recovery action based on the guardrail feedback. The action is executed only if it passes both the safety and task-progress checks. We evaluate AgentRecover on AgentDojo and Agent Security Bench (ASB): 1,749 attack variants of 107 user tasks across 14 application scenarios. Across five diverse LLMs on AgentDojo, AgentRecover reduces the mean attack success rate (ASR) from 27.54% to 0.00% and increases mean under-attack utility from 57.74% to 87.40%. Local comparisons against guardrail and recovery baselines show substantial gains. These results show that AgentRecover can preserve agent safety while substantially improving utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.