The Free-Recipe Limit: Every Recipe Effect Measures Which Premise Broke
Abstract
Post-training pipelines must choose which skills to train together, whether to run one blocked stage per skill or interleave them, and in what order, and a large literature reports that these choices matter. Whether any of them can be decided depends on two quantities that are never reported together: how far apart the best and worst recipes can be, and how finely a re-execution of the whole protocol can tell two recipes apart. Where the second exceeds the first, no quantity of search converts into a decision, and the signature is not an absence of winners but winners that do not survive re-running. We measure both sides for supervised fine-tuning on competition mathematics, pricing the resolution from nested variance components over repeated executions of the protocol rather than from seed spread alone. Across six skill pairs the largest order contrast on any single corpus is 0.062 while one execution resolves only 0.081 — a ratio of 0.76 that is an upper bound, since its numerator is the maximum of a search — and re-running the winner reverses its sign. Most of that noise is cheap to remove: at four samples per problem, evaluation sampling rather than training stochasticity carries 0.955 of the variance of the order contrast itself, so quadrupling the samples costs minutes and beats a fourth training seed costing hours. One departure clears the wall: split a corpus so its halves answer under incompatible but equally correct conventions, and arrangement stops moving capability and starts moving which convention the model commits to, by two orders of magnitude over its same-convention control, across four scales in two pretraining families. Recipe ablations should therefore publish their re-execution resolution beside their effect; a public scorecard grades all 30 preregistered claims, of which 7 failed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.