acceptodds
Under review as a conference paper at ICLR 2027

Auditing Agent Evaluators: Repair Coverage, Evidence, and Explanation Provenance

Abstract

Completion scores guide agent comparisons, yet overall detection can obscure a semantic judge's added coverage beyond existing repairs. We audit public τ-bench scorers and an adapted Tool-Veritas semantic component on fixed executions, linking coverage, judge inputs and votes. For selected failures, we inspect supplied evidence and test controlled input changes. The uniform author reapplication of authorization rules to 236 retail executions finds 88 local violations still accepted by the revised database scorer. DeepSeek Flash detects 28 of these from raw events versus seventeen from extracted features. Among 117 database-accepted executions satisfying the local requirements, the same conditions reject three and zero, respectively. On Flash, source-linked facts detect more violations overall than raw events but fewer residual violations. Eighteen of 2,024 semantic outputs display a reason from a vote opposing the criterion majority. Selecting a supporting vote fixes this source-selection defect without changing labels. These results motivate tracking additional detections, newly rejected locally valid behavior, and each displayed reason's source.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.