acceptodds
Under review as a conference paper at ICLR 2027

Ciallo: Correlation-Aware Multi-Objective Alignment for Large Language Models

Abstract

Aligning large language models (LLMs) with diverse human preferences requires balancing multiple objectives that may be compatible or conflicting in different scenarios. However, existing multi-objective alignment methods often optimize different objectives independently, ignoring their correlated preference patterns and conflicting constraints. This oversight prevents the model from reusing shared information across compatible objectives and leveraging contrastive signals from conflicting ones, thereby reducing training efficiency. Motivated by this consideration, we propose Ciallo, a framework that first conducts separated training to learn an initial soft prompt representation for each alignment objective, and then performs joint training under mixed preference weights, where a dynamically tracked Fisher-gradient affinity matrix constrains objective relationships and guides Group Relative Policy Optimization (GRPO) through a Fisher Gram-matching regularizer. Experiments on Reddit Summary with Qwen3-0.6B and Helpful Assistant with Qwen3-4B show that Ciallo consistently achieves macro-averaged gains of and reward points over Reward Soups.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.