Budgeted Discovery of Latent Reliability Failures in Large Language Models
Abstract
Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but opera- tionally important failures. A prompt that produces no failures in a small sample may therefore appear reliable despite having a nonzero latent probability of failure under repeated inference. We formulate LLM reliability evaluation as a budget-constrained discovery prob- lem in which each prompt is associated with an unknown per-generation failure probability. We propose a budgeted discovery framework that first performs shal- low evaluation across the prompt set and then uses trial-level failure outcomes together with prompt-derived representations to learn a feature-based ranking of failure propensity. The resulting scores prioritize prompts with zero observed shallow failures for deeper evaluation, concentrating the deep-evaluation budget where hidden failures are more likely to be discovered. We evaluate this approach on AIRBench and StrongREJECT across multiple model and system-prompt conditions. The central empirical test asks whether models fit without access to deep-evaluation outcomes can rank prompts with zero observed shallow failures according to their likelihood of producing failures under deeper evaluation. On AIRBench, the highest-ranked 10% of unresolved prompts achieves 2.54× hidden-failure lift for Qwen 2.5 7B and 1.87× for Gemma 3n E4B, recovering 25.4% and 18.7% of subsequently observed hidden failures, respectively, compared with 10% expected under random allocation. Semantic- neighborhood and feature-ablation analyses further show that this predictive sig- nal can be recovered from multiple representations of prompt content and relationships. These results support treating the reliability of an LLM for a given prompt as a latent stochastic property and show that prompts with zero observed failures under shallow evaluation can nevertheless differ in their latent failure probabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.