TangentMix: Recovering Specialized Language Models through State-Conditioned Data Allocation
Abstract
Specialized language models can acquire strong domain capabilities while developing uneven regressions across others. In on-policy distillation (OPD), controlled interventions show that source exposure can harm its matching capability and the same mixture can reverse its effect as the specialist evolves. These failures expose a mismatch between common data-value signals and the multi-capability effect of the complete update received by the current learner. We propose TangentMix, a state-conditioned data-allocation method that jointly adapts source exposure and complete recovery–preservation update selection. TangentMix evaluates candidate updates from the current model and optimizer state, commits the selected update, and uses independently measured post-selection utility to update future source allocation. Across four recovery settings spanning 1.7B and 4B students and a broad capability suite covering world knowledge, compositional reasoning, multilingual comprehension, mathematics, code generation, and instruction following, TangentMix improves 92.9% of evaluated capability axes over their specialist initialization. On the three 4B specialists, it closes 79.1% of the aggregate recovery gap and achieves the highest Macro-7 in every profile; across all four settings, it raises the average Macro-7 by 4.2% relative to standard OPD. It further improves competition-level AIME24/25 reasoning by 1.88 points over OPD and achieves 2× data efficiency over fixed-mixture MOPD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.