Lost in Formulation: Disentangling and Measuring Prompt Sensitivity
Abstract
Language model outputs vary with both the phrasing of a prompt (prompt formulation) and the amount of information it provides about the intended question (prompt specificity). Separating these sources of sensitivity helps users decide how to rewrite a prompt, but existing prompt sensitivity metrics conflate them. We vary prompt formulation and prompt specificity systematically on questions from the AmbigQA dataset, collect repeated answers from three open-weight models, and compute ten existing metrics on these answers. A principal component analysis separates the metrics into three components, which we call correctness, formulation dependence, and output dispersion. While existing metrics capture correctness and output dispersion well, none captures formulation dependence without also counting decoding noise, the variation among repeated answers to the same prompt. We therefore introduce , an intraclass correlation of correctness across formulations that subtracts decoding noise. Although it does not enter the decomposition, tracks the formulation-dependence component more closely than existing metrics, also on held-out AmbigQA questions. Varying specificity and formulation confirms that the three components respond to different changes. Disambiguating a question raises correctness and lowers output dispersion but produces no detectable change in , whereas more varied formulations tend to raise but not the other two. Correctness, formulation dependence, and output dispersion are thus distinct parts of prompt sensitivity, and measures formulation dependence separately from decoding noise.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.