The Hidden Cost of Prompt Formatting: Measuring Representational Sensitivity in Language Models
Abstract
Prompt formatting is often treated as a superficial implementation detail, yet meaning-preserving format changes can alter model behavior. We ask whether such changes also leave a systematic footprint inside the model before decoding, and whether this internal movement is associated with output instability. We introduce an activation-level audit for prompt-format sensitivity using matched one-shot prompt pairs in which the task instruction, demonstration, input, and answer semantics are fixed while exactly one structural formatting property is perturbed. Across 53 Super-NaturalInstructions tasks and four open-weight 7B-scale models—Falcon, LLaMA-2, Mistral, and Qwen-2.5—we analyze 1,739,400 atomic interventions per model spanning separator, enumeration, spacing, and casing edits. We find a stable hierarchy of representational sensitivity: separator edits induce the largest hidden-state shifts, enumeration edits form a consistent second tier, and spacing and casing produce smaller effects. We further evaluate activation–output coupling over 5,573,004 eligible base–perturbation pairs spanning 168 task–model strata. Controlled activation-distance coefficients are positive in 158/168 strata, while 167/168 within-stratum Spearman correlations are positive, with median , showing that larger activation shifts are generally associated with larger changes in semantic-label output distributions. Finally, we introduce PFSS, a compact 25-term interpretable predictor of activation movement from prompt-format metadata alone, achieving task-holdout raw-scale scores of 0.831, 0.773, 0.852, and 0.821. These results establish prompt formatting as a measurable component of the representational input space and activation movement as an informative, though heterogeneous, intermediate signal of output instability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.