LEMMA: Learned Model Merging for Text-to-Image Generation via Multi-Objective Optimization
Abstract
Merging specialist text-to-image models into a unified generator can consolidate diverse generative capabilities while minimizing deployment overhead. However, existing methods rely on static, global heuristics that ignore layer-specific functionality and fail to balance multiple objectives. To address this, we propose LEMMA, a novel framework that formulates text-to-image model merging as a multi-objective optimization problem. LEMMA adaptively combines specialist task vectors through layer-tailored merging coefficients that are learned directly on the target objectives, while the base and specialist models remain frozen. We demonstrate LEMMA on two foundation architectures (SDXL and SD3) across two multi-objective setups, by merging reward specialists under multi-objective Direct Preference Optimization (DPO), and by merging style specialists under multi-objective Supervised Fine-Tuning (SFT). Extensive experiments show that LEMMA consistently outperforms prior merging baselines across both setups. Importantly, LEMMA often surpasses dedicated single-task specialists on their own target metrics, improving all objectives simultaneously rather than trading one capability for another. These gains are generalized to unobserved evaluation metrics and are corroborated by a user study with human raters, all while introducing no additional inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.