When Richer Preferences Do Not Help: Rubric Redundancy, Pair Selection, and Training Exposure in Scientific QA
Abstract
Fine-grained rubrics and dense preference pairs appear to offer richer supervision for aligning scientific-domain large language models (LLMs), but nominal supervision richness need not translate into non-redundant usable preference information. We study this question on 4,000 titanium-alloy reverse-design questions, each associated with four scored candidate answers and task-specific rubric vectors. In the audited single-checkpoint matrix, clean-start binary and six-pair graded DPO both display approximately 92.6 mean generated-answer score at one-decimal precision. Rubric dimensions are strongly correlated for several reasoning types, and at q50 rubric-absolute and scalar-gap pair selection overlap substantially (Jaccard 0.943), although the scalar score is itself aggregated from the rubric dimensions. Across three training seeds, rubric-aware q50 filtering does not establish a reliable advantage over scalar or random selection. Moreover, a seed-42 truncated full-data probe reaches 92.85 after approximately 9.7k preference-pair exposures, compared with 92.31 after a full pass over 19.2k pairs, showing that one-epoch filtered-versus-full comparisons change training exposure together with pair composition. Together, these results suggest that, in this controlled setting, nominal supervision richness can be a poor proxy for non-redundant reference information; preference-signal efficiency should therefore be analyzed through feedback redundancy, pair composition, and training exposure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.