acceptodds
Under review as a conference paper at ICLR 2027

Closed-Form Least-Squares Forecasting for Fast Diffusion Transformers

Abstract

Diffusion transformers spend most of their inference budget re-evaluating a network whose outputs change slowly along the sampling trajectory. Cache-based accelerators exploit this redundancy by skipping selected steps, and their quality is governed by two levers—which steps to compute and how to fill in the skipped ones—that have mostly been studied in isolation. We build a framework that exposes both at once and use it as a controlled attribution study. Our predictor treats the sequence of whole-transformer outputs as the observable of a discrete-time system and forecasts a skipped output as a linear functional of previously computed outputs, fitted either per image by global least squares or, once for all images, by closed-form calibration on a handful of prompts requiring no gradient training. The attribution result is unambiguous and survives every control we apply: step selection dominates predictor design, and per-image adaptive gating is worse than committing to a schedule even at a lower skip rate. Repeating the whole protocol on SD3.5-Medium, a B MMDiT-X model with flow matching and a different text stack, shows the same ordering and a better operating point: wall-clock at CLIP below the uncached model with of its sharpness, and CLIP at or above the uncached model at a -step budget and at px. On PixArt- the method reaches – at alignment matching the uncached model on two benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.