acceptodds
Under review as a conference paper at ICLR 2027

When Richer Preferences Do Not Help: Rubric Redundancy, Pair Selection, and Training Exposure in Scientific QA

Abstract

Fine-grained rubrics and dense preference pairs appear to offer richer supervision for aligning scientific-domain large language models (LLMs), but nominal supervision richness need not translate into non-redundant usable preference information. We study this question on 4,000 titanium-alloy reverse-design questions, each associated with four scored candidate answers and task-specific rubric vectors. In the audited single-checkpoint matrix, clean-start binary and six-pair graded DPO both display approximately 92.6 mean generated-answer score at one-decimal precision. Rubric dimensions are strongly correlated for several reasoning types, and at q50 rubric-absolute and scalar-gap pair selection overlap substantially (Jaccard 0.943), although the scalar score is itself aggregated from the rubric dimensions. Across three training seeds, rubric-aware q50 filtering does not establish a reliable advantage over scalar or random selection. Moreover, a seed-42 truncated full-data probe reaches 92.85 after approximately 9.7k preference-pair exposures, compared with 92.31 after a full pass over 19.2k pairs, showing that one-epoch filtered-versus-full comparisons change training exposure together with pair composition. Together, these results suggest that, in this controlled setting, nominal supervision richness can be a poor proxy for non-redundant reference information; preference-signal efficiency should therefore be analyzed through feedback redundancy, pair composition, and training exposure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.