DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
Abstract
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving previous capabilities, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristics or adaptive rules that cannot explicitly enforce preservation of such capabilities. We propose DynaMiCS, a dynamic mixture optimizer that casts multi-domain fine-tuning as a constrained optimization problem. At each update, DynaMiCS performs short domain-specific probing runs to estimate a slope matrix of local cross-domain effects, capturing how training on each fine-tuning dataset affects each evaluation domain. These estimates are then used to compute mixture weights through optimization over the probability simplex, with the objective of improving target-domain performance while keeping constrained-domain metrics within a specified tolerance of reference levels. Because these effects are measured by finite differences rather than gradients, targets and constraints need not be differentiable, or present in the fine-tuning data, and can be specified directly as benchmark accuracies. Across scenarios with varying numbers of target and constrained domains, and with loss- or accuracy-based objectives, DynaMiCS achieves stronger target-domain improvements and higher constraint satisfaction than static, dynamic, similarity-based and probing-based alternatives, without a reference model, per-example scoring, or manually tuned weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.