acceptodds
Under review as a conference paper at ICLR 2027

The Evidentiary Control Protocol: A Preregistered Governance Layer for Auditable AI Evaluation

Abstract

Scores reported by AI benchmarking frameworks do not, by themselves, establish how an observed answer was produced. When the question of interest is whether an outcome was obtained under a specified experimental condition — rather than through retrieval, memorization, model-mediated inference, or answer-bearing artifacts in the evaluation environment — the score is not the evidence. The evidence is a controlled chain from case construction to outcome classification. We present the Evidentiary Control Protocol (ECP), a preregistered, leakage-controlled governance layer that fixes a ten-stage lifecycle — case authoring, structural and novelty review, environmental knowledge-boundary verification, registration, conditioned execution, raw-evidence preservation, independent audit, mechanical outcome classification, descriptive statistical analysis, and representation-bias disclosure — and enforces one central separation discipline: execution validity, evidence integrity, protocol-defined outcome, and scientific interpretation are four distinct claims, recorded separately and never merged. ECP does not measure a model's internal state and does not adjudicate cognitive claims. The paper demonstrates the protocol operating across an evidence portfolio that now extends beyond the original five lines (the T1 pilot; the M1 pilot across six hash-frozen systems; the REASON-2 campaign; the REASON-3 defect-preservation/repair/re-execution chain; and the M2 comparative recording campaign) to an independent audit programme that applies the same discipline to externally published claims under bounded channels: A14: Commercial structured-output audit of Together AI — schema validity 24/24 = 1.000 under the registered configuration A15: Adversarial claim-injection audit — detection 1.00 on levels L1–L3, silent acceptance 0.00 A16: Adaptive adversarial red team — 141 submissions across four valid runs, 0 accepted, with a documented localization-metric boundary and final classification INCONCLUSIVE A17: Structured-output reliability audit of OpenAI's published claim — 60/60 schema-valid, verdict CLAIM_HOLDS_STRONG, bounded to the tested corpus, snapshot, and gateway channel A18: Bounded directional audit of Anthropic's published prompt-injection robustness claim — mechanical injection-success 3/30 = 0.100 on Sonnet 4.6 versus 14/30 = 0.467 on Sonnet 5, non-overlapping 95% Wilson intervals — a directional difference opposite to the published claim's direction under the tested conditions We are explicit about what this establishes and what it does not. The portfolio establishes the protocol's operation — registrations held, evidence preserved, adverse findings kept intact — and demonstrates that ECP-style independent audits can examine third-party commercial claims with bounded channel disclosure. It does not validate ECP as a scientific instrument for measuring reasoning, does not replicate any vendor's internal evaluation, and does not convert any bounded finding into a general claim. Every quantitative statement is bound to a sealed artifact in the accompanying claim ledger.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.