Worlds Before Words? A Causal Boundary Audit of Procedural Pre-Pretraining
Abstract
Does procedural pre-pretraining transfer useful computation, or do apparent gains depend on budget and training boundaries? We study this distinction with paired initializations and downstream minibatches in an auditable small-decoder regime. Across 365 frozen trajectories, typed-Dyck warmup improves over downstream-only scratch in 14/15 pilot pairs by learning-curve area and 15/15 by final bits per byte. Stronger controls change the interpretation. Continuous natural training and retention of natural embedding/head parameters each win all 30 confirmation-plus-scale cells. Dyck beats exact within-document shuffling in all 30, but a first-order Markov surrogate strictly beats Dyck in 29/30. These directions persist at both tested widths. A separate checkpoint diagnostic shows that Dyck models did learn bounded nonadjacent closing-type information that their Markov controls did not. Learned procedural information therefore need not yield a downstream advantage in this protocol. A reduced-scale bridge to a positive reference regime changes the task to ordered UNION and lowers the procedural allocation to about 0.3% of training model-tokens. UNION then beats all three declared comparators on final validation loss in two originally completed seed blocks and a separately disclosed post-outcome resource completion. This is not an exact reference replication or an isolated test of allocation. The evidence supports a regime-specific boundary, not general failure of procedural pre-pretraining. Additive gains, learned procedural capability, and fixed-budget downstream value are distinct claims that require separate controls.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.