acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Text-First Pretraining for Audio Language Models

Abstract

How should a fixed pretraining budget be allocated when extending a language model from text to audio? Large Audio Language Models (LALMs) are often built by starting from a pretrained text-only language model and then adding audio tokens through continued pretraining. This text-first strategy reuses an already capable linguistic backbone, but it also commits part of the total training budget to a text-only objective before the model begins joint text–audio learning. Determining whether this sequential allocation is more effective than learning both modalities from the beginning requires a controlled comparison. We therefore compare text-first continued pretraining (CPT) with joint text–audio pretraining from scratch (SPT) under matched end-to-end training budgets. We conduct a controlled sweep over training budgets, text–audio mixture ratios, and model scales while holding the decoder-only architecture and data sources fixed across CPT and SPT. Across the evaluated settings, SPT generally achieves a better balance between text and audio performance than CPT. Our results suggest that CPT can be inefficient because part of the budget is spent optimizing a text-only objective before training shifts to the joint text–audio objective, while text replay further reduces the budget available for audio-side training. SPT's advantage, however, does not appear to arise from simple positive cross-modal transfer, as the two objectives can still interfere. Instead, the results suggest that allowing both objectives to shape the model throughout training contributes to a better balance between text and audio capabilities. This trend extends to a Mixture-of-Experts architecture and a speech-generation objective, and persists after supervised fine-tuning on speech question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.