acceptodds
Under review as a conference paper at ICLR 2027

Syntactic Shadowing: Why Length Penalties Do Not Fix Preference Optimization

Abstract

Preference optimization is supposed to teach a language model what people prefer, but a training pair exposes more than meaning. The resulting shortfall has been attributed to length bias and answered with length penalties and length normalization, as if controlling length were controlling non-semantic bias. These remedies have not fixed preference optimization: models trained with them align meaning no better than models trained without them. We ask whether the reason is a confound in the training signal itself. A longer response is rarely just a longer version of a shorter one; it carries different sentence structure. Length and syntax are confounded in natural text, so a model that appears to prefer long responses may in fact prefer the syntax that long responses tend to have. To make this measurable, we use meaning-preserving rewrites in which each step changes exactly one property, yielding a length gap measured at fixed syntax and a syntactic gap measured at fixed length, both read against the scale of the discrepancy between two rewrites of the same response produced under the same specification. Building on this, we propose a framework that decomposes the alignment gap of a preference pair into a length, a syntactic and a semantic component, together with a construction that rewrites the rejected response into token-exact counterfactuals of the chosen one and computes the corresponding gaps. Applying this to 25 models spanning ten state-of-the-art methods, and to 16 models of a second backbone, reveals a consistent pattern. Whether trained or not, and whichever method is used, directional semantic preference does not track the length channel, and the syntactic channel beneath it stays where it was; what does improve it, in a controlled comparison on identical data, is training on surface-matched counterfactual pairs. We call this syntactic shadowing: the syntactic component of surface bias sits in the shadow of length, where length control does not reach it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.