acceptodds
Under review as a conference paper at ICLR 2027

Convex Multi-Target Dataset Selection for Post-Training via Portfolio Optimization

Abstract

Post-training multiple target capabilities under a fixed auxiliary-data budget is harder than single-target selection because one auxiliary mixture must account for transfer across potentially conflicting targets. Inspired by single-target data selection with Kernel Mean Matching (KMM), we view this problem through the lens of Markowitz portfolio optimization: auxiliary datasets are assets, target-auxiliary alignment is return, and the similarity matrix among auxiliary datasets is the risk in the capital asset pricing model. This view naturally extends to multiple targets by letting a user-specified preference vector define the aggregate return across target capabilities, while the redundancy penalty among auxiliary datasets remains shared. Based on this formulation, we introduce Multi-Target KMM (MT-KMM), which uses linear scalarization to compute continuous dataset scores, then post-processes them into a common auxiliary training mixture for multiple targets. Motivated by Adam's first update from zero moments, we represent each dataset by its normalized signed gradient at a frozen reference model, which coincides with Adam's normalized first-step features from zero moments in the zero-stabilizer limit. The gradients are computed only once, and the resulting selection can be reused to train models up to 70B. Theoretically, we show that finite-preview MT-KMM valuations converge to their population counterparts at the rate for raw gradients and, under a linear near-zero condition, for SignGD, where is the size of the preview gradients. Empirically, MT-KMM achieves the highest mean target accuracy among the compared methods at every tested Llama scale and auxiliary-data budget for SFT, with gains extending to Qwen models from 0.6B to 32B and to RL post-training. By varying the target preference vector across different target capabilities, we show that MT-KMM can adapt the selected auxiliary mixture to different downstream priorities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.