CertiLoop: Auditing Learned Agent Artifacts under Benefit–Risk Contracts
Abstract
An average benchmark gain does not establish that a learned agent artifact should be reused under explicit risk constraints. We introduce CertiLoop, an audit protocol that compares each complete frozen artifact against its comparator, fixes risk strata from initial task information before execution, nests benefit–risk contracts (benefit-only, whole-scope, and stratum-guarded) on shared intervals so that each refutation is attributable to a specific requirement, and reports every decision as positive, negative, or unresolved. We audit 30 artifacts from ACE, GEPA, and ReasoningBank on fixed AppWorld and -Retail rosters, using 8,192 independent paired executions per artifact as reference evidence. Development selection agrees with every resolved benefit label, yet the reference refutes five of the 25 selected artifacts (20%) on risk grounds, three of them only once the stratum conditions are added. All eight judge configurations we evaluate recommend at least one reference-negative artifact, including six that retain every reference-positive artifact, and task-level prediction accuracy does not order the reliability of the resulting recommendation sets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.