acceptodds
Under review as a conference paper at ICLR 2027

TangentMix: Recovering Specialized Language Models through State-Conditioned Data Allocation

Abstract

Specialized language models can acquire strong domain capabilities while developing uneven regressions across others. In on-policy distillation (OPD), controlled interventions show that source exposure can harm its matching capability and the same mixture can reverse its effect as the specialist evolves. These failures expose a mismatch between common data-value signals and the multi-capability effect of the complete update received by the current learner. We propose TangentMix, a state-conditioned data-allocation method that jointly adapts source exposure and complete recovery–preservation update selection. TangentMix evaluates candidate updates from the current model and optimizer state, commits the selected update, and uses independently measured post-selection utility to update future source allocation. Across four recovery settings spanning 1.7B and 4B students and a broad capability suite covering world knowledge, compositional reasoning, multilingual comprehension, mathematics, code generation, and instruction following, TangentMix improves 92.9% of evaluated capability axes over their specialist initialization. On the three 4B specialists, it closes 79.1% of the aggregate recovery gap and achieves the highest Macro-7 in every profile; across all four settings, it raises the average Macro-7 by 4.2% relative to standard OPD. It further improves competition-level AIME24/25 reasoning by 1.88 points over OPD and achieves 2× data efficiency over fixed-mixture MOPD.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.