From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents
Abstract
LLM-based guardrails typically safeguard agents by assessing proposed agent actions in context before execution, producing safety signals such as binary allow/deny decisions, risk categories, and explanatory rationales for potential policy violations. However, agent risks often arise when otherwise benign tasks are contaminated by untrusted external content, unsafe instructions, or tool misuse. Existing guardrails often flag the entire task uniformly as unsafe, thereby blocking the threat but sacrificing the benign part. Moreover, existing work largely evaluates guardrails in isolation, leaving unclear whether their interventions effectively guide agents toward safer downstream behavior. To address this, we introduce **TRIAD** (Tripartite Response for Iterative Agent Guardrailing), an agent-guardrail framework that produces safety feedback at each planning step and provides it as additional context for the agent to remediate misaligned action plans before downstream execution. We deploy Tri-Guard within TRIAD as a guardrail model fine-tuned on a self-curated training dataset to generate a detailed safety analysis of the proposed action plan along with a three-way decision: *proceed*, *refuse*, or *update*. This feedback is incorporated into the agent's context to guide subsequent planning. When malicious prompts redirect the agent toward unintended goals, Tri-Guard guides the agent to remediate the proposed plan, avoid misalignment, and preserve the original goal where possible. Across ASB and AgentHarm, TRIAD reduces the average attack success rate to 10.42% while offering the most favorable safety–utility trade-off among the evaluated baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.