acceptodds
Under review as a conference paper at ICLR 2027

MOSAIC: More LLMs, Less Cost through Collaborative Reasoning for Compiler Optimizations

Abstract

LLM-guided compiler optimization has recently shown promise. However, current LLM-guided optimizers query a single LLM at every search step, so their cost grows with both that model’s price and the search budget. A smaller model is cheaper but degrades the search: in our experiments, a recent single-LLM optimizer reaches 10.2% (6.0%) lower speedup on GPU (CPU) with gpt-5-mini than with GPT-5.2, which gives up the output quality that motivates using LLMs in the first place. Agentic multi-LLM frameworks add controllers, concurrent calls, and shared memory, whose overhead can offset the savings of cheaper models. This paper shows that the optimization search tree itself can coordinate heterogeneous LLMs: the tree already records which proposals improved the program, and we can augment the tree to record which model made them. That evidence can then be used to decide which model should propose next. To this end, we introduce MOSAIC, in which each Monte Carlo tree search (MCTS) node records a program and the LLM that expands it, and the active LLM proposes both a compiler transformation and the LLM for the next step. We also contribute an LLM-aware UCT that favors smaller models and a course-alteration rule that invokes the largest model only after repeated small-model regressions. Against that single-LLM optimizer run with GPT-5.2 alone and the same number of search iterations and tree parameters, on an NVIDIA GPU (Intel CPU), MOSAIC with eight LLMs improves the geometric-mean speedup on five kernels by 7.2% (14.4%) and the end-to-end Llama-3-8B speedup by 1.61× (1.41×), while reducing kernel compilation time by 1.95× (1.74×) and API cost by 4.47× (4.32×), with GPT-5.2 making only 23.1% (23.9%) of LLM calls. With the same eight models, random or round-robin model choice improves over the single-LLM baseline by only 2.5% on CPU, showing that the gains come from letting the LLMs choose the next model rather than from mixing the pool at random.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.