Prompt Conditions Can Change Measured Intervention Effects in LLM Evaluation
Abstract
Prompted language model evaluations often report the effect of an intervention under a single prompt condition. We test whether that reported effect changes when the surrounding prompt condition changes. On MMLU-Pro, a silent self-check increases exact-format accuracy for Gemma by under a standard prompt condition, under a masked condition, and under a token-matched metadata condition. A separate, pre-specified 300-task replication gives gains of , , and . The self-check gain is therefore larger under Standard than under Masked, with a simultaneous 95% confidence interval of . DeepSeek shows little observed difference across the same prompt conditions and Llama is inconclusive. The result also depends on scoring. On the replication set, an automated semantic answer scorer gives a Standard-minus-Masked difference of , while a separate constrained-decoding evaluation gives . Supporting experiments further show that average accuracy can remain almost unchanged even when many individual outcomes flip. These findings motivate reporting intervention gains under the prompt and scoring conditions that define them, rather than treating those choices as incidental details. Code: https://anonymous.4open.science/r/PCCMIE_ICLR27-BCA1/README.md
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.