ChiEngMixBench: When “Model Preference” Depends on How You Score It
Abstract
Technical concepts can often be expressed either with English terms or their Chinese equivalents. A common way to ask which form a language model favors is to compare the likelihoods the model assigns to the two forms. Yet such “preference” is not directly observed from the model; it is induced by the scoring rule used to measure it. Whether we score the whole sentence or the tokens overlapping the changed text, and whether we sum or average token-level scores, can change the resulting judgment. We introduce ChiEngMixBench, a controlled testbed for studying this issue through paired Chinese–English terminology substitutions, and compare several natural likelihood-based definitions of preference. We find an easily overlooked pattern: scoring definitions can agree on the majority direction while disagreeing substantially about the winners for individual pairs. Thus, agreement on the majority direction alone does not establish which individual items or subgroups behave consistently across scoring definitions; robustness must be checked at the level of the claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.