Piggybacking Training on Inference Systems
Abstract
Training a model on the inference system that serves it would enable private on-device fine-tuning and turn idle inference fleet capacity into training capacity without separate training infrastructure. Inference hardware, however, often lacks the memory and interconnect needed for backpropagation. Zeroth-order optimization, such as evolution strategies (ES), uses only forward passes, but evaluating candidate-specific perturbations adds weight-update traffic or separate correction costs to batched inference. We address this problem with NoisyGEMM, a matrix-multiplication kernel that evaluates perturbed models without materializing their weight perturbations. Candidates use compact coefficient vectors over shared Hadamard bases, whose corrections are fused into the base GEMM. We integrate NoisyGEMM into vLLM's linear layers and run training through its existing scheduler, KV-cache manager, and sampler. On an A100, perturbed linear layers run at 1.1×–1.3× the latency of unperturbed cuBLAS GEMM. Using a separable CMA-ES variant, reinforcement learning of Qwen3 models from 1.7B to 30B-A3B improves on the base model on all six reasoning benchmarks and outperforms GRPO on 17 of 18 model–benchmark pairs with the same maximum completion length of 16k tokens. Compared to the state-of-the-art ES framework EGGROLL, NoisyGEMM hosting CMA-ES achieves higher accuracy on AIME 2025 and lower latency when scaling the population size and adapter rank. Finally, we demonstrate piggybacking CMA-ES training on a serving engine that replays one day of BurstGPT live traffic. We will open-source our code upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.