Rethinking Targeted Data Selection for Instruction Tuning: Balancing Target Specialization and Cross-Task Utility
Abstract
Instruction data selection aims to identify a compact subset of training examples that improves downstream performance under a fixed data budget. Existing target-specific methods primarily optimize relevance to a designated target task, but improvements on the target may coincide with cross-task degradation, while task-balanced methods typically treat multiple tasks symmetrically without explicitly prioritizing the target. We therefore rethink targeted data selection for instruction tuning as an asymmetric objective: improving target specialization while preserving cross-task utility. To address this objective, we propose STAR (Task-Shared and Target-Specific Activation-based Ranking), which uses cross-task FFN activation statistics from a frozen LLM to construct two complementary neuron sets. Task-Shared neurons exhibit consistently high and relatively balanced activation across reference tasks, while Target-Specific neurons exhibit stronger activation for the designated target than for the remaining reference tasks. These neuron sets define separate activation subspaces, where STAR measures candidate alignment with all reference-task prototypes and the target prototype, respectively, and combines the two scores for selection. We evaluate STAR across five target tasks and four OOD benchmarks under selection budgets of 5% and 1%. Under the 5% budget, STAR achieves the highest Target Avg. of 62.80 and Non-Target Avg. of 60.21, along with the best aligned OOD performance and cross-target OOD transfer among target-specific methods. STAR remains effective under the 1% budget and across backbone models. These results demonstrate that STAR improves target specialization while preserving cross-task utility under limited budgets. The code is available at https://anonymous.4open.science/r/STAR-42E7/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.