Conjugate Compilation: Self-Improving Agentic Co-Compilation for Speculative Decoding
Abstract
As LLM adoption surges, scaling model serving requires not only adding more GPUs but also unlocking the substantial capacity stranded in those already deployed. Speculative decoding brings two models into the serving process: a small draft model proposes tokens that a larger target model verifies. Yet compiling these models individually leaves their shared execution outside the compiler’s optimization boundary. We therefore ask whether co-locating the draft and target models on the same GPU and co-compiling their execution can reclaim this stranded capacity. We introduce CONJUGATE COMPILATION, which reframes speculative decoding as a cross-model compilation problem and jointly optimizes the draft and target models within their shared execution rather than treating each in isolation. One of our key insights is that cross-model compilation should be informed by the measured performance of the co-located shared execution it produces: an optimization that improves a kernel in isolation need not improve the models when they execute together. To expand cross-model opportunities, we devise operand layouts that interleave blocks from both models within common tiles, co-locating GEMM and non-GEMM work while preserving their mathematics. We realize CONJUGATE COMPILATION through a self-improving agentic search system in which an LLM edits cross-model optimization plans using the compiler and GPU execution as its environment. Each proposed plan is compiled and evaluated through its co-located shared execution, whose measured score, together with kernel-level verdicts and compiler feedback, informs subsequent proposals, while Monte Carlo tree search accumulates rewards to guide exploration. The system also expands its search space and tuning budget when progress stalls and retains tuned kernels and verdicts for reuse across runs. Across six single-GPU speculative-decoding configurations on NVIDIA B300 and RTX 5090 GPUs, CONJUGATE COMPILATION achieves a geometric-mean per-round speedup of 1.37× over conventional execution. Interleaved kernels raise this speedup to 1.45×, reaching 1.59× in the best case. CONJUGATE COMPILATION also improves GPU utilization on average by 1.46× on RTX 5090 and 1.50× on B300. These results demonstrate the potential of cross-model conjugate compilation guided by measurements of the co-located shared execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.