Training is Serving: Unifying Inference and Adaptation for Multimodal Extension
Abstract
Extending large language models (LLMs) to multiple modalities is a key step toward general-purpose multimodal intelligence. As foundation models scale, conventional fine-tuning increasingly assumes homogeneous clusters and pays model-scale communication across ranks. We present Training is Serving (TiS), an API-driven compute-sharing framework that decouples the heavy computation of a frozen LLM from lightweight multimodal adaptation. Specifically, we enable the LLM to support training and return gradients in the form of an API. An edge client trains only its modality encoder and adapter, sends embeddings to the service, and receives gradients; raw data and trainable parameters remain at the edge, while the LLM weights remain at the cloud server. TiS decouples edge-side data parallel training from provider-side model parallel serving, which has four core advantages: (1) compute sharing, where edge clients can perform training by accessing the API without acquiring server control; (2) hardware decoupling, where heterogeneous accelerators are unified behind a standard API backend; (3) low-bandwidth communication, enabling multi-node MLLM training over only 1 Gbps links; and (4) data locality, where raw training data remains on the edge while the service receives only embeddings and labels. Using shared compute, we extend Gemma-4-31B-it to the speech modality and support 70 languages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.