acceptodds
Under review as a conference paper at ICLR 2027

JointValid: Prediction-Powered Multi-Risk Certification of Large Language Models

Abstract

Large language models are evaluated across heterogeneous capabilities and risks, yet multi-dimensional scorecards rarely attach an error rate to the deployment claim that every required risk lies below its threshold. We formulate this decision as multi-risk certification with abundant scores from imperfect black-box judges and a shared human-annotation budget. Within each task, prediction-powered inference uses human scores to correct a larger judge-scored sample. Across tasks, an intersection–union test asymptotically controls erroneous global certification, with an explicit normal-approximation remainder and without multiplicity correction, cross-task independence, or calibrated judges. We derive a certification-aware allocation from a finite-judge-pool approximation to joint power. Under safe alternatives, the objective is strictly concave, and a scalar Lagrange dual yields its global solution. An independent pilot estimates task margins and residual uncertainty; a confirmatory reserve prevents task starvation, while a fractional-count Beta working posterior stabilizes allocation with limited pilot data. Experiments on three human-rated suites scored by open-weight judges show when adaptive allocation improves power over uniform allocation and when limited task heterogeneity or pilot information constrains its gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.