acceptodds
Under review as a conference paper at ICLR 2027

When Reasoning Traces Overrun the Canvas of Diffusion LLMs

Abstract

A diffusion language model (dLLM) decodes into a fixed-length canvas, and supervised fine-tuning (SFT) on teacher traces shapes how much it writes before answering. We show that the SFT corpus affects such a student largely through overrun: reaching the end of the canvas without an answer, which we measure directly on saved generations of LLaDA-8B students. Holding teacher, problems and correctness fixed, prompting the teacher for terse rather than detailed traces raises GSM8K accuracy by 15.8 points at a 512-token canvas, and most of the gap (14.3 points) is attributable to overrun. A statistic computed before training, the share of training responses longer than the canvas, orders student overrun across corpora. A second check exposes a defect in the standard d1 pipeline: its 4,096-token SFT cap truncates about 90% of s1K examples and leaves only about one in ten with its answer, which teaches the student not to terminate. On LLaDA-8B-Base, where s1K SFT is no better than no SFT, replacing it with a corpus that fits the canvas improves accuracy by 8.8 GSM8K points, attributable to overrun, and by 13.7 MATH500 points, about half attributable to truncation. On LLaDA-8B-Instruct, which rarely overruns in-domain, s1K remains better in-domain; but on Countdown the s1K-trained student overruns on most problems, before and after identical GSM8K RL, and this accounts for the budget-fitting pipeline’s advantage there.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.