acceptodds
Under review as a conference paper at ICLR 2027

Prompt Conditions Can Change Measured Intervention Effects in LLM Evaluation

Abstract

Prompted language model evaluations often report the effect of an intervention under a single prompt condition. We test whether that reported effect changes when the surrounding prompt condition changes. On MMLU-Pro, a silent self-check increases exact-format accuracy for Gemma by under a standard prompt condition, under a masked condition, and under a token-matched metadata condition. A separate, pre-specified 300-task replication gives gains of , , and . The self-check gain is therefore larger under Standard than under Masked, with a simultaneous 95% confidence interval of . DeepSeek shows little observed difference across the same prompt conditions and Llama is inconclusive. The result also depends on scoring. On the replication set, an automated semantic answer scorer gives a Standard-minus-Masked difference of , while a separate constrained-decoding evaluation gives . Supporting experiments further show that average accuracy can remain almost unchanged even when many individual outcomes flip. These findings motivate reporting intervention gains under the prompt and scoring conditions that define them, rather than treating those choices as incidental details. Code: https://anonymous.4open.science/r/PCCMIE_ICLR27-BCA1/README.md

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.