Pre-Pre-Training Before Pixels and Genes
Abstract
Foundation models trained on distinct modalities are hypothesized to converge toward similar representations. We turn this observation around and ask whether some universal prior can be learned from synthetic data before exposure to natural data and how the learned priors transfer across different domains. We investigate this question through synthetic pre-pre-training, in which transformers are first trained on data generated from structured processes without any target-domain semantics followed by domain-specific pre-training. We treat synthetic pre-pre-training as a design space, varying the synthetic substrate (neural cellular automata or Dyck formal language), training objective (masked reconstruction or autoregressive prediction), and target domain (vision and genomics). Across three vision benchmarks and ten genomics tasks, we observe that synthetic pre-pre-training generally improves downstream performance, but that its effectiveness depends on our studied three design factors. In vision, masked reconstruction consistently outperforms autoregressive prediction for both pre-pre-training sources. Masked 2D NCA pre-pre-training achieves the highest accuracy on all three benchmarks, improving upon the randomly initialized and trained baseline by on average. In genomics, both synthetic sources transfer effectively, with autoregressive pre-pre-training on Dyck sequences outperforming the randomly initialized baseline by approximately 6 points on NTv3 mean score and 2.6% on the Genomics Classification tasks. These results suggest that synthetic pre-pre-training is a promising way to instill general computational priors before domain-specific pre-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.