acceptodds
Under review as a conference paper at ICLR 2027

CertAgent: Conformal Risk Control for Agent Runtime Fragility Certification

Abstract

LLM agents extend language models from text generation to interactive task solving, but their deployment can fail through fragile runtime behaviors arising from reasoning or action errors. Existing mitigation methods improve empirical robustness, but they do not provide formal risk guarantees over the entire Agent Runtime. In this paper, we propose CertAgent, a framework for Agent Runtime Fragility Certification. We first formulate Agent Runtime Fragility as system level risk over the full agent runtime trajectory, while using stage specific failure modes to guide runtime control. CertAgent instantiates a Verifier Subagent to detect unreliable runtime proposals, a Gate Module to control runtime conservativeness through a configuration parameter , and a reflection mechanism to revise rejected proposals. To provide deployment time guarantees, we apply Conformal Risk Control to certify Agent Runtime Fragility, enabling users to either obtain a certified runtime risk bound for a fixed configuration or identify certified runtime configurations under a target risk level. We further introduce policy optimization to reduce system level runtime risk and improve the resulting certification. Experiments on representative agent benchmarks show that CertAgent improves runtime reliability, supports practical risk utility tradeoffs, and yields more effective certification. Our code is available at https://anonymous.4open.science/r/CertAgent.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.