acceptodds
Under review as a conference paper at ICLR 2027

NuAlign: A Numerical Distance Perception Alignment for Large Language Models as Evaluators

Abstract

Large language models (LLMs) have shown huge potential as automatic evaluators, but the standard token-based likelihood objectives don't explicitly encode the numerical distance between scores. Previous research introduced numerical supervision through expected score regression or token-level numerical loss; however, the former can't distinguish distributions with the same average, and the latter doesn't differentiate the digit positions. To address this, we propose NuAlign, a numerical distance perception alignment for LLMs as evaluators that directly uses the full score deviations reconstructed during training to update the likelihood of the true score sequences. NuAlign consists of two stages. In Stage I, we learn to evaluate semantic and scoring boundaries through parameter-efficient supervised fine-tuning, thereby establishing a task-aware autoregressive scoring distribution. In Stage II, we utilize numerical scoring deviations to modulate token-level likelihood optimization, introducing numerical distance sensitivity while preserving the native autoregressive scoring distribution. We conducted experiments across multiple evaluation datasets and model families, achieving an MAE of 0.28 on the Mohler dataset (out of a maximum score of 5). Our code and data are available at: https://anonymous.4open.science/r/I-LOVE_ICLR-7561

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.