acceptodds
Under review as a conference paper at ICLR 2027

ReMix: Efficient Diffusion Serving with Adaptive Reuse and Model Mixture

Abstract

Serving text-to-image diffusion models online is expensive, because each request runs tens of sequential denoising steps through a large network. Mixture-of-models serving reduces this cost by caching images that a large model has already generated and repainting them with a small model for similar prompts. Existing methods select the repainting strength by fixed hand-crafted rules and repaint the entire image, which miscalibrates the strength and overwrites content that was already correct. We propose **ReMix**, an efficient diffusion serving framework that reuses cached images and generates with a mixture of models. **ReMix** profiles the denoising trajectory of each cached generation to select a per-request repainting start step and repaints under a graduated change map that expands from the high-change regions to the full image, leaving remaining pixels at the fidelity of the cached image. On COCO and DiffusionDB, **ReMix** cut the mean serving latency by 52.3% and 26.3% against full large-model generation, and achieves the state-of-the-art CLIP, FID, IS, and PickScore among all serving methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.