How Evaluation Choices Affect MAUVE Scores and Generator Rankings
Abstract
Since its introduction, has been widely adopted for evaluating open-ended text generation, serving both to compare language models and to design and select decoding strategies. Because it compares generated and human-written text in a learned feature space, its scores depend on evaluation choices such as the representation model, the evaluated text length, and the quantization settings. These choices are known to shift scores, but less is known about whether they also change which generation model is preferred. We study this question across seven generation models, three datasets, and 180 evaluation settings, separating variation due to evaluation choices from variation due to decoding configurations. For identical generations, evaluation choices produce a wider score spread than decoding changes, and they alter model rankings even after aggregating over decoding strategies: two settings that differ only in the representation model select the same top-ranked model in 30% of comparisons. Among the choices examined, the representation model has the largest effect on ranking stability, and it also affects how closely rankings agree with human preference judgments. These findings suggest that interpreting a MAUVE comparison requires considering its evaluation configuration and the sensitivity of model rankings to that configuration, and show that human evaluation can identify representation models whose rankings align more closely with human preferences, providing practical guidance for choosing among them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.