HSIR: Hidden-State Intrinsic Rewarding for Multilingual Alignment
Abstract
Large language models (LLMs) exhibit strong yet uneven multilingual capabilities, motivating alignment methods that improve cross-lingual consistency without costly external evaluators. Existing approaches often derive rewards from translation models, multilingual encoders, or stronger language models, incurring additional inference costs and introducing a mismatch between external reward spaces and the policy model. We propose HSIR, a Hidden-State Intrinsic Rewarding framework for multilingual alignment. HSIR constructs an intrinsic reward directly from the policy model's hidden representations by measuring the similarity between generated responses and reference answers, and incorporates it into Group Relative Policy Optimization (GRPO). This enables multilingual alignment without an additional reward model or translation evaluator. Experiments on multilingual open-ended generation show that HSIR consistently outperforms supervised fine-tuning and alternative semantic, cross-lingual, and translation-based rewards. On Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct, HSIR achieves All-avg win rates of 49.02% and 66.32%, respectively, corresponding to improvements of 21.58 and 20.31 percentage points over the original models. These results demonstrate that hidden representations provide an effective intrinsic reward for efficient multilingual alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.