Population-Based Diversity for Reinforcement Learning in LLMs via Sequential Multi-Adapter Optimization
Abstract
Reinforcement learning (RL) fine-tuning of large language models (LLMs) tends to collapse the policy onto a few strategies. This induces a severe quality-diversity (QD) trade-off: while single-policy performance improves, solution coverage (e.g., pass@) shrinks, leaving the model brittle when its preferred strategy fails. While other methods may opt to maintain diverse modes of behavior within a single policy, we instead use a framework that sequentially builds a population of LoRA adapters over a shared frozen backbone, optimizing each policy for task performance while rewarding distribution-level diversity against previously learned policies conditioned on sufficient task competence. Unlike response-level diversity objectives, our method compares the cross-policy probability distributions, avoiding surface-level similarity metrics or additional learned diversity models. On long-horizon BabyAI tasks, our method improves held-out pass@ and is more robust to prompt-based distribution shift than independently trained policies. We also demonstrate that our approach seamlessly transfers to large-scale jailbreak red-teaming
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.