Certify, Refute, or Abstain: Claim-Aware Policy Auditing in Multi-Agent Reinforcement Learning
Abstract
High task return does not establish that a policy satisfies a declared protocol or deviation-gap claim. Claim-Aware Protocol Auditing (CAPA) distinguishes the evidence needed to certify a claim from the evidence sufficient to refute it: an upper bound supports certification, a lower bound supports refutation, and insufficient evidence requires abstention. We formulate unconditional false-decision and false-selection guarantees and show that passive rollouts cannot improve a worst-case gap bound on an indistinguishable bounded-game pair. A constructed soft leader–follower example separates return ranking from protocol-map conformity: the return-10 candidate violates the map tolerance, whereas the selected conforming candidate has return 3.956. Synthetic boundary calibration exposes undercoverage of Gaussian plug-in bounds; matched Student-t bounds require iid Gaussian gains. On fixed learned MPE candidates, a seeded neural-deviation audit reports all 45 probes with paired returns and explicit Gaussian-approximation limits. No learned candidate has an unrestricted upper-bound certificate. The resulting evaluation contract makes policy performance, evidence of a violation, and the stronger evidence needed for certification separate objects of assessment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.