acceptodds
Under review as a conference paper at ICLR 2027

Success Has a Shape: TailPass@k for Repeated Sampling Evaluation

Abstract

Repeated-sampling evaluation measures language models over multiple non-deterministic attempts on each task, but scalar summaries such as accuracy, , , or a single thresholded repeated-sampling score can conflate systems with different operational behavior. A model that occasionally finds a correct answer can have high discovery but low repeatability, while another model may have lower single-sample accuracy but much stronger consistency across attempts. We characterize failure modes of commonly used evaluation metrics and address information loss and uncertainty through a discovery–stability tail profile: for each task, reporting budget , and threshold , the profile is the probability that the model produces at least correct responses among attempts. We place a Dirichlet–categorical model on each task's latent outcome distribution, with binary exact match recovered as the Beta–Bernoulli special case, and propagate this uncertainty to dataset-level profiles (across tasks), yielding closed-form posterior means and covariances, together with posterior Monte Carlo credible intervals. The resulting uncertainty-aware report is TailPass@. The framework recovers discovery, , stability, and single-sample accuracy as special cases and shows that accuracy is the area under the profile. Utilities over this profile combine its threshold probabilities into scalar scores that express preferences over discovery and repeatability, with uncertainty propagated to each score (e.g., power-moment utilities shift emphasis from discovery toward repeatability as their parameter increases). Exact enumeration of response patterns shows that endpoint metrics and uniform-threshold area can merge response patterns with the same total correctness but different discovery–stability structure, while the full profile exposes these distinctions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.