acceptodds
Under review as a conference paper at ICLR 2027

CertiLoop: Auditing Learned Agent Artifacts under Benefit–Risk Contracts

Abstract

An average benchmark gain does not establish that a learned agent artifact should be reused under explicit risk constraints. We introduce CertiLoop, an audit protocol that compares each complete frozen artifact against its comparator, fixes risk strata from initial task information before execution, nests benefit–risk contracts (benefit-only, whole-scope, and stratum-guarded) on shared intervals so that each refutation is attributable to a specific requirement, and reports every decision as positive, negative, or unresolved. We audit 30 artifacts from ACE, GEPA, and ReasoningBank on fixed AppWorld and -Retail rosters, using 8,192 independent paired executions per artifact as reference evidence. Development selection agrees with every resolved benefit label, yet the reference refutes five of the 25 selected artifacts (20%) on risk grounds, three of them only once the stratum conditions are added. All eight judge configurations we evaluate recommend at least one reference-negative artifact, including six that retain every reference-positive artifact, and task-level prediction accuracy does not order the reliability of the resulting recommendation sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.