Off-Menu Preferences Diagnose Menu Sensitivity in Forced-Choice Evaluation
Abstract
Forced-choice benchmarks score a language model only on the answers they list (the menu), yet a model can score an unlisted answer above every listed one. We ask whether such off-menu preferences indicate sensitivity to the menu or an incorrect answer. A gold-free off-menu ratio compares the best-scoring candidate in a specified excluded set with the top listed answer; a controlled swap replaces the weakest distractor with that candidate while retaining the gold answer. Across four open-weight checkpoints, controlled cloze tasks, and four established benchmarks, the menu can matter more than the model: on 941 single-token CommonsenseQA items, the swap lowers Qwen2.5-1.5B's accuracy from 66.0% to 21.6%, a 44.4-point drop, more than twice the largest accuracy gap between checkpoints (20.3 points). Because the inserted candidate is never the gold answer, the drop equals the proportion of items on which it displaces a correct answer; displacement rate times overall accuracy tracks drops of 3 to 44 points across three checkpoints and four benchmarks (correlation 0.998). At equal coverage, the margin between the top two listed answers is the stronger selector of correct answers (81.1% vs. 68.4% on CommonsenseQA; likewise on ARC and OpenBookQA). The ratio instead flags risk under the altered menu: used in labelled calibration at 50% coverage, it lowers the proportion of Qwen2.5-1.5B's samples from that menu that miss the gold answer from .652 to .610. Evaluations should report the menu, the excluded-candidate search, and the calibration target alongside accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.