acceptodds
Under review as a conference paper at ICLR 2027

Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models

Abstract

Diffusion large language models (dLLMs) support parallel decoding and bidirectional context, but competitive dLLMs have 8–16B parameters, which limits deployment. Knowledge distillation is a natural remedy, yet existing dLLM methods only reduce inference steps: the distilled model is faster but no smaller. Distillation for model compression, routine for autoregressive (AR) LLMs, has not been studied for dLLMs. Because teacher and student differ in backbone, attention, and tokenizer, compressing a dLLM is a cross-architecture problem with three obstacles: timestep-dependent teacher reliability, context scarcity from heavy masking, and tokenizer/vocabulary mismatch. We present TIDE, a regime-adaptive framework whose components are matched to the teacher–student gap: TIDAL (timestep-aware scheduling) and CompDemo (context enrichment) for shared-tokenizer pairs, and Reverse CALM (noise-filtering chunk-level alignment) for cross-tokenizer pairs. TIDE distills 8B-dense and 16B-MoE teachers, one per run, into a 0.6B student that needs 11–22× less memory than the teachers. Naive distillation does not help: it causes negative transfer for the shared-tokenizer pair and gives no gain for the cross-tokenizer pair. TIDE improves on it in both pipelines (+3.00 and +1.95 on the eight-benchmark average) and ends above the non-distilled student (+0.88 and +1.53), most clearly on GSM8K (+6.68 in the cross-tokenizer pipeline) and less clearly on code. The reported scores of a same-size AR model remain higher (40.9 vs. 34.20 average). To our knowledge, this is the first evidence that a large dLLM can be compressed across heterogeneous architectures. Code is available at https://anonymous.4open.science/r/tide-5E36/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.