acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Learning for Fair LLM Fine-Tuning from Imbalanced Training Signals

Abstract

Fine-tuning large language models (LLMs) using logged reward data can inadvertently lead to unfair performance disparities across prompt groups, particularly when training data is imbalanced or groups differ in inherent difficulty. In the context of personalized recipe generation, we observe that naive fine-tuning to optimize towards the expected reward severely degrades performance for underrepresented or challenging prompts (e.g., specific cuisines or rare ingredients). This motivates the need for more principled approaches to fine-tuning LLMs in the presence of biased distributions and varying task difficulties. We propose a novel adversarial offline policy optimization method for LLM fine-tuning, aiming to improve worst-case group performance, thereby addressing performance disparities. Our adversarial method alternates between reweighting training samples and updating the policy, which supports a wide range of fairness goals such as enforcing fairness for specific vulnerable groups, achieving balanced performance across all possible groups, and optimizing for uniform relative group-level improvements compared to a reference policy. We empirically validate our proposed approach through controlled experiments in recipe generation and content summarization. Our method consistently improves worst-case group performance, narrowing across-group performance gaps observed in baseline methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.