acceptodds
Under review as a conference paper at ICLR 2027

Latent Rubrics: Identifiability and Semantic Transfer in Preference Models

Abstract

Can reusable evaluation criteria be learned without updating a language-model backbone? We study parameter-efficient semantic grounding on frozen representations: our Qwen3-32B main study trains only a 15,363-parameter rubric head, with no backbone updates. The central challenge is identification, not parameter scale. Preference likelihood is invariant to changes of latent basis; exact named interventions identify directions up to positive scales under a restricted linear model. A synthetic positive control raises axis recovery from 0.730 to 0.996 with one intervention per criterion, while preference accuracy changes little. On real text, a human-audited filter passes 19/20 fresh blind pairs, but locally valid edits alone do not ensure transfer: Silver-200 obtains 61.40% TEST rubric composition versus 72.57% for DEV-aligned preference-only training. A subsequent 10,240-parameter semantic-ranking head, using additional human labels and retaining the frozen backbone, reaches 72.67% DEV composition: 3.81 points above the natural-contrast intervention variant, but not a confirmed improvement over aligned BT. This demonstrates a lightweight route to improving development-set semantic metrics, not established TEST superiority. Our framework combines restricted identification theory, auditable supervision, and controlled diagnostics to distinguish parameter-efficient learning from the stronger requirement of reusable semantic axes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.