acceptodds
Under review as a conference paper at ICLR 2027

AdmitScope: Compiling Evidence-Grounded Leaderboards into Audit-Qualified Claims

Abstract

Evidence-grounded leaderboards guide scientific comparison, but their score margins are trustworthy only when abstention policies, evaluated populations, and cited documents support the same conclusion. AdmitScope makes that validity condition executable. Given a proposed row comparison, it compiles task scope, evidence access, coverage policy, metric, and uncertainty into a claim; activates obligations for paired margins, population comparability, document grounding, held-out stability, and independent review; and issues a direct, transferred, or surrogate-only license with an inspectable evidence trace. We instantiate AdmitScope on 10,261 eligibility, endpoint, and safety decisions grounded in public protocol and safety records, with 2,586 risk-based-monitoring cases as a transformed-input stress setting. On the 2,565-case three-family frozen-test panel, retrieval adds 0.117–0.123 EP across matched GPT-4.1, Claude-3.5, and Llama-3.1 70B backbones; calibrated selection adds 0.051–0.084 EP and reduces overconfidence by 8.3–8.9 points at controlled coverage. GPT-4.1 + RAG + selective reaches EP 0.932 and EF1 0.897, while open Qwen-2.5-72B reaches EP 0.881 [0.861,0.900] and EF1 0.851 [0.829,0.872]. Three blinded clinicians reviewing 800 covered outputs assign the leading system DocEP/DocER of 0.91/0.86 and flag 90.1% of high-overlap wrong-evidence pairs, compared with 47.9% under span overlap. A 1,600-output row-stratified audit, 4,640 matched-case judgments, a 384-case independent clinician evaluation, two independent assembly paths, a 600-case source-native cross-rubric evaluation, post-cutoff records, cue weakening, and five-seed reruns execute the remaining obligations. The resulting 48-claim frontier grants direct support to 36 claims, transfers 6, and assigns 6 to surrogate support while exposing 2 reversals hidden by raw rank; all 18/18 within-backbone regime claims survive. AdmitScope turns leaderboard validity into a reproducible, decision-bearing output.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.