Consensus-Guided Preference Optimization for Multi-Source Language Model Alignment
Abstract
Large language models often possess complementary capabilities, yet combining those capabilities in one deployable model remains difficult when source models differ in architecture, specialization, and probability calibration. Most fusion approaches rely on parameter merging or supervised fine-tuning. Existing preference-stage alternatives depend critically on the quality of constructed chosen-rejected pairs. We introduce Consensus-Guided Preference Optimization (CGPO), a framework for multi-source language model alignment from unpaired preference supervision. CGPO turns source-model sequence probabilities into a confidence-calibrated consensus that directly guides each pointwise preference update. The consensus supplies a completion-specific optimization baseline, while pointwise verifier feedback identifies desirable and undesirable responses. This design transfers complementary capabilities from heterogeneous sources without parameter merging, vocabulary alignment, or artificially constructed preference pairs. Comprehensive experiments on seven benchmarks spanning mathematics, code generation, reasoning, and instruction following show that CGPO outperforms existing model-fusion and preference-optimization baselines. For Phi-3.5, CGPO raises the average performance across these benchmarks from 58.1 to 63.3. Our code is available at https://anonymous.4open.science/r/CGPO-6501/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.