acceptodds
Under review as a conference paper at ICLR 2027

Ermine: Neural Differential Scaling Laws for Diffusion Language Models

Abstract

Scaling laws prescribe how to spend compute; for diffusion language models (DLMs), they should also prescribe *when*. Existing compute-optimal recipes for DLMs inherit the Chinchilla scaling law, which allocates capacity and data uniformly across the diffusion process. We observe, however, that DLM loss varies substantially with the diffusion timestep, leaving compute underutilized under Chinchilla-style allocation. We introduce *Ermine*, a differential scaling law that characterizes the instantaneous loss via time-dependent capacity and data densities. We derive its optimal solution under a fixed compute budget and realize it with quantile-based stratified sampling and a sparse Mixture-of-Experts (MoE) architecture. Ermine accelerates Chinchilla-optimal training by , improves compute efficiency over dense DLMs by , and narrows the compute-optimal gap between DLMs and autoregressive models (ARMs) from to . On OpenWebText, our model attains a state-of-the-art validation perplexity of among DLMs, and at FLOPs it outperforms existing ARMs and DLMs on five zero-shot commonsense reasoning benchmarks using only -% of their training FLOPs. Our analysis further reveals that Ermine generalizes better by avoiding excessive compute in time regions where scalability diminishes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.