Puro-2B: Poor Lab’s Qwen2-1.5B Trained on RTX 5090 within $5090
Abstract
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over 700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models differ in token budgets and selected recipe variants. Our best model is trained for less than 4.4K, less than $5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, our end-to-end case study shows that the curriculum-based pretraining advantage persists under matched supervised fine-tuning, with higher performance on mathematical and general-purpose tasks. Access to the full pretraining pipeline rather than model weights alone enables such controlled studies. We release the full Puro-2B training recipe, including data, code, and model weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.