acceptodds
Under review as a conference paper at ICLR 2027

Mixed precision inference for Diffusion language models

Abstract

Diffusion Language Models (DLMs) have drawn increasing attention as an alternative to classical Large Language Models architecture due to their potential to significantly increase inference speed. However, they lack in practice some key optimizations. One such feature is mixed-precision inference, i.e., using a quantized or a full precision model at different steps of the generation process. This hybrid design enables one to benefit from the accelerated inference of quantized models while maintaining an output quality similar to that of original floating-point precision. As existing approaches are not applicable to this setup, we propose here the first mixed precision framework specifically targeted for DLMs. By modeling the quantization error, we are able to design a training-free, adaptive mechanism to successfully exploit the speed of a quantized model and the precision of a regular one. Our results show we can achieve a 66% speed increase without significant degradation of the model's performance. We also demonstrate this can be combined with other acceleration systems to further speed up the sampling process.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.