NAPO: Neuron Activation-based Preference Optimization for Multilingual Reasoning
Abstract
Large language models exhibit uneven reasoning abilities across languages due to imbalances in multilingual training data. Recently, preference optimization has substantially improved multilingual reasoning performance. However, existing methods typically derive preference signals either from external translation models or from coarse-grained internal signals, such as hidden-state representations. These signals may entangle reasoning-relevant information with language-specific or surface-form noise, resulting in inaccurate preference estimation and ultimately limiting improvements in multilingual reasoning. To address this issue, we propose **N**euron **A**ctivation-based **P**reference **O**ptimization (**NAPO**), a framework that derives cross-lingual preference signals from internal fine-grained neuron activations. Specifically, we first identify reasoning-relevant neurons shared across languages. We then introduce a position-weighted alignment reward computed in the identified neuron space to construct pairwise preference data. Finally, we iteratively optimize the policy within the standard direct preference optimization framework. We conduct extensive experiments on three popular multilingual mathematical reasoning benchmarks across 10 languages. The results show that NAPO consistently outperforms multiple strong preference optimization-based baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.