Inference-Native Zeroth-Order Optimization for LLMs
Abstract
Zeroth-order (ZO) methods train large language models using only forward passes, yet common ZO execution paths perform substantial work beyond what the algorithm itself requires. To remove this extra execution overhead, we present Infer-ZO, which separates the evaluations required by the algorithm from how they are executed, allowing them to reuse existing inference-engine optimizations. After removing inherited execution work, a complete Infer-ZO step on Qwen3-14B adds only 1.16% wall-clock time over matched inference, bringing ZO execution close to its fundamental inference workload. With Infer-ZO, ZO evaluations run at near-inference cost with frozen base weights and can share GPU batches with ordinary inference requests. In co-serving experiments, foreground and background throughput remain within 1–3% of the corresponding background-request baseline. Across 15 Qwen3, Llama, and OPT models, Infer-ZO achieves 2.10×–9.94× end-to-end step speedups over released LoZO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.