acceptodds
Under review as a conference paper at ICLR 2027

The Moving Frontier of Native Speech–Language Pretraining

Abstract

Multimodal large models are typically built through a late-fusion recipe: language models and modality-specific components are pretrained separately, then connected and jointly trained. This approach builds on advances in unimodal pretraining and has become a standard route to multimodal capabilities. Recently, frontier systems have begun to explore a different recipe, learning perception and language jointly from scratch. These developments suggest that native multimodal pretraining is a viable alternative and motivate a closer examination of the principles governing its scaling. In this work, we study this paradigm through continuous speech and text, using speech–language modeling as a concrete setting in which to investigate how model capacity, training compute, and modality-specific supervision interact. Across over 140 completed configurations spanning 149M–2.72B parameters, we demonstrate that randomly initialized models can jointly acquire speech recognition and text-modeling capabilities. We characterize how scale, capacity placement, and supervision shape speech and text learning, and use boundary and initialization controls to probe these relationships. A scaling law connects these empirical findings to loss prediction across configurations. Together, these results provide an empirical foundation for native speech–language pretraining and a quantitative guide to allocating its training resources.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.