acceptodds
Under review as a conference paper at ICLR 2027

ASR Data is All You Need to Build Instruction-Following SpeechLLMs

Abstract

SpeechLLMs can follow instructions like text LLMs by connecting a pretrained speech encoder to an LLM through an adapter. Current approaches typically rely on a two-stage training procedure including fine-tuning on instruction-following (IF) speech data, which is expensive to collect. In this work, we investigate whether IF speech data can be avoided by using Sequence-Level Knowledge Distillation (SeqKD) from a pretrained LLM operating on automatic speech recognition (ASR) data, exploring both prompt-conditioned SeqKD, where the same prompt is provided to the student and teacher, and prompt-free SeqKD, where no prompt is provided. With experiments on three tasks (ASR, speech translation, and spoken question answering) and four benchmarks including text instructions and outputs in languages unseen during training, we show that SeqKD effectively transfers the LLM's linguistic and IF capabilities but suffers from poor ASR performance. We address this limitation with a single-stage training strategy combining prompt-free SeqKD with limited ASR and prompt-conditioned data. Our approach outperforms a cascade of ASR and LLM and a two-stage approach, and remains competitive across tasks with SpeechLLMs trained on substantially larger amounts of data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.