When Formatting Determines Rankings: Separating Semantic Accuracy, Compliance, and Strict Accuracy
Abstract
Language-model evaluations can conflate correctness, format adherence, and parser recognition. Across five models, MATH-500 and a balanced 13-domain MMLU-Pro subset, we vary final serialization while preserving the reasoning instruction and report extracted-answer accuracy, terminal-grammar compliance, and strict interface accuracy. The original 24,750 generations comprise a 17,250-response main experiment and a paired token-budget study. Paired intervals and multiplicity correction identify 12 strict-accuracy and 5 compliance rank reversals, but no reliable extracted-answer accuracy reversal. Nine original strict reversals persist under a wrapper-tolerant policy. A targeted 2,000-response, same-session intervention removes angle brackets around the answer placeholder, eliminating the Qwen32–DeepSeek-V3.2 Final/JSON strict point-order reversal on both tested subsets. A further study yields 3,000 outputs on held-out questions: complete prospective MMLU-Pro results support the prespecified model-by-format template interaction, while MATH results after an additional recovery request constitute a protocol-deviation sensitivity analysis, not overall prospective confirmation. Template effects are heterogeneous, including substantial accuracy declines in some conditions, and do not establish a universal repair. These findings demonstrate sensitivity of specified model–protocol snapshots, not invariant model deficiencies or latent capability ordering. Evaluations should report extracted-answer accuracy, compliance, and strict accuracy separately, alongside prompts, extraction rules, and token budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.