: Meta-Evaluating Metrics for Multi-lingual Style-Personalized Generation
Abstract
Meta-evaluation of metrics and LLM-judges for text generation is primarily done using metric-human correlations, even though prior work has shown that correlation does not give an appropriate estimate of a metric's evaluation capability, and can be unstable across settings such as different datasets or languages. We propose a regression-based meta-evaluation of metrics that learns on calibrated metric outputs to predict human ratings. We ground this measurement relative to a trivial constant-mean baseline and the human agreements, to provide a comparable measurement of metric capability across evaluation settings. To test our meta-evaluation framework, we developed a large-scale multi-lingual style personalized text generation (SPTG) benchmark, collecting multiple annotations over 10 topologically diverse languages and four evaluation dimensions of varying subjectivity. We meta-evaluated 32 metrics from three style evaluation paradigms (string-based, embedding-based, and LLM-judges) and their ensembles over these 40 evaluation settings. Our analysis revealed that even simple strong-based and embedding-based measurements can match strong LLM-judges when calibrated to human ratings. We also identified characteristics of human annotations that determine the performance of metrics, as well as when a metric ensemble outperforms its individual constituents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.