Rethinking Multi-Objective Resolution in Federated LLM Alignment
Abstract
Aligning large language models (LLMs) requires balancing potentially conflicting objectives, such as helpfulness and harmlessness. Federated learning can support alignment when preference data are distributed across clients, but multi-objective alignment introduces a distinct conflict-resolution challenge. Server-side conflict resolution may require transmitting an objective-specific gradient for every client and objective, whereas independent client-side resolution can produce inconsistent objective weights because of stochastic estimation, local model drift, and client heterogeneity. We call the resulting inconsistency *multi-objective disagreement drift*. We propose **FIRM** (**F**ederated **I**n-client **R**egularized **M**ulti-objective LLM alignment), a communication-efficient client-centric framework that resolves objective conflicts locally and communicates one model update per client. FIRM regularizes each client's multiple-gradient descent (MGDA) subproblem, reducing the sensitivity of its objective weights to perturbations in estimated gradients. For a federated multi-objective actor-critic abstraction, we derive a finite-time bound on the average Pareto-stationarity gap that separates standard federated client drift from the additional disagreement induced by client-side conflict resolution. The bound establishes convergence to a residual neighborhood of Pareto stationarity and characterizes how this disagreement depends on regularization strength, batch size, local updating, and client heterogeneity. Experiments on federated LLM alignment show that FIRM reduces client objective-weight disagreement and improves the observed reward trade-off relative to the evaluated baselines. Additional experiments demonstrate stable behavior under heterogeneous reward models, stronger non-IID partitions, larger client populations, partial participation, additional objectives, and larger model scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.