acceptodds
Under review as a conference paper at ICLR 2027

Using eye-tracking to understand and estimate reliability of preference labels

Abstract

Large Language Model alignment methods like DPO rely on availability of high quality preference annotations, where annotators are asked to choose which of typically two responses they prefer. Often, it is assumed that these human-annotated labels are reliable in that they represent preferences well. However, these annotations may not reflect underlying preferences due to factors like labeling errors, task difficulty, biases introduced by the experimental setup, etc. While examining the impact of noisy labels on downstream tasks has gained attention in the literature, less emphasis has been placed in understanding where the noise comes from—how annotators complete the annotation tasks. In this paper, we use eye-tracking to examine how preference annotation is done and collect an eye-tracking dataset EyeTracking-HS2-800 using Apple Vision Pro (including 5 annotations for approximately 800 tasks from the HelpSteer2–Preference dataset completed by 100 participants). Our results point to positional biases in the distribution of visual attention over the samples, and two distinct strategies for completing the annotation tasks emerge: "readers" display an increasing task completion time and “scanners” a relatively constant task completion time with the length of the shown sample. We then use these observations to modulate weight of individual human-annotated preference labels so that more carefully annotated samples get a higher weight (measured by proportion of fixated text and reading speed). By increasing the weight of the "reader" annotations, we recover, and on average even outperform, the performance of all HelpSteer2-Preference labels despite having significantly fewer annotations, outperforming all other training configurations. Based on our results, we provide guidelines for conducting and reporting preference annotation studies. In particular, we suggest incorporating reading speed, or time spent per task, as a suitable alternative to improve LLM alignment in the case of annotator disagreement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.