SynthCode: Not All Synthetic Code Is Worth the Same Tokens
Abstract
Synthetic data now underpins code pretraining, yet the decisions that produce it are made largely by intuition: which transformation to apply to a source file, whether to train on the generated artifact alone or pack it with its source, and whether a fixed token budget is better spent on new documents or on new views of the same ones. Because these choices are rarely compared at equal cost, their relative value is unknown. We introduce SynthCode, a 5.1M-document corpus distilled from 6.4M Python commits with gpt-oss-120b, together with a token-matched protocol for isolating the value of each generation decision. The protocol covers sixteen recipes: single-transformation distillation, five context-packing formats, equal-token mixtures, and a three-way source-overlap study that prices a second view of one document against a first view of two. Every recipe is trained for the same number of tokens under identical hyperparameters, so differences are attributable to the data rather than to compute or tuning. We then train the best-performing recipe into six base models spanning two families and 4B–31B parameters (Qwen3-4B/8B/14B, Gemma4-12B/26B/31B). We release the corpus, the generation and filtering pipeline, and all checkpoints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.