Morphology-Consistent Preference Optimization for Indic Language Models
Abstract
Preference optimization has become the standard practice for aligning large language models, yet every dominant objective scores whole responses or raw subword tokens. The mismatch is acute in morphologically rich Indic languages, where one grammatical unit can decide rejection while the rest of the response stays correct. Case suffixes, agreement and tense-aspect-mood markers, honorific forms, and script-sensitive vowel signs all carry that weight. Subword tokens do not fill that role either as one token may conflate a stem with a suffix, and one morpheme may split across tokens. A token-level objective therefore spreads its signal past the unit that decided the preference. We introduce MorphDPO, which moves the unit of alignment from a flat token sequence to a latent morpheme lattice. The resulting reward replaces the sequence-level Direct Preference Optimization (DPO) log-ratio with a normalized morpheme-lattice reward. The new reward is tokenizer-agnostic, marginalizes over uncertain analyses, and reduces exactly to DPO under uniform weights. The underlying tilted-policy log-ratio keeps morphology weighting a valid change of policy geometry rather than an unnormalized token reweighting. Furthermore, we build three components, namely counterfactual partial-order preferences from rule-generated minimal pairs, a paradigm-neighborhood objective enforcing consistency across valid inflections, and a group-distributionally-robust objective over language-script-morphology groups. Across eight Indic languages on the Indic-specialized Sarvam-1 2B, MorphDPO attains the best composite accuracy () and win-rate () among preference objectives, with the highest morphology minimal-pair accuracy and a group gap tied with DPO for the smallest. The same lead holds on Llama-3.2 3B and Qwen3 1.7B, proving that morphology rather than the token is the right unit of preference alignment for morphologically rich languages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.