acceptodds
Under review as a conference paper at ICLR 2027

Everything in Moderation: Interior Per-Domain Coverage Optima and Gap-Preserving Fixed-Budget Alignment in Multi-Domain Mid-Training

Abstract

Mid-training data composition is usually set by availability, not design. We test whether downstream alignment repairs this on Qwen3-8B-Base, using a mid-training-only 4B arm to probe scale dependence across five rule-disjoint KOR-Bench domains. We sweep 30 allocations on the simplex with five seeds each. The internal mid-training token share leaves a between-domain gap structure that alignment shifts but does not close. First, every domain's coverage marginal has an interior optimum: the moderate 10–40 percent band is best for all five domains, supported by a family-wise quadratic-interiority permutation test at P approximately 0.010, though the boundaries are not independent of the fitted peaks and the decline above the band resolves in only one domain. Coverage is collinear with repetition, but a share-by-pool-fraction experiment quadruples the epoch count while accuracy moves at most 1 percentage point, so repetition does not carry the optimum on the tested domain. Second, the gaps survive: compensatory SFT bridges 0/240 pairs at 5 percentage points, and a budget-feasible permutation null still places it below chance (P = 0.005), as does a uniform control. Third, zero coverage collapses a domain below its no-mid-training baseline (Counterfactual 86.0 percent to 45.6 percent), although not all domains do so (Cipher gains 8.8 percentage points); SFT removes the collapse, and two ladders attribute it to the domain's absence (8B; the 4B arm does not reproduce it). On the six withheld allocations, the mid-training ranking survives alignment (r = +0.975) and the fitted curves add no measurable increment (r = +0.967; both at n = 6), while a mixture baseline lands within 1 percentage point of θ*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.