LLM Benchmark Means Can Be Blind to the Tail
Abstract
LLM benchmark failure rates are often read as properties of model weights, but they are also properties of the decoder used to sample outputs. We show that this matters most in the tail. On independently selected SecurityEval prompts where -detected weaknesses are rare under an untruncated baseline, lowering the temperature to reduces prompt-weighted medium-severity rates by roughly one third to one half across three models, with rate ratios between and . On the full benchmark, the corresponding ratio for Qwen2.5-1.5B is , with a 95% paired-prompt bootstrap interval of . All code-weakness rates are conditional on parseable output, and parseability is reported separately. This contrast motivates treating the decoder as part of the estimand. We study when samples collected under one decoder can estimate a rate under another. For configurations that share the same model logits, a teacher-forced pass over stored completions gives exact likelihood ratios for target decoders, without additional generation. The transfer is valid only when the audit support contains the target's event-relevant support. When it does, we use a per-prefix Rényi-2 quantity computed from the stored logits as an empirical predictor of transfer cost, and show that prefix weighting can improve efficiency for prefix-measurable events. The experiments also show a limit of reuse: measured effective sample size and other sample-based diagnostics can miss poorly sampled mass, and do not by themselves support a family-wide certificate. We recommend reporting the decoding configuration with every failure rate, and checking support, statistical cost, and parseability before transferring conclusions across decoders.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.