Small Loss, Poor Generator: Where Consistency Tuning Should Train for One-Step Generation
Abstract
Training a one-step image generator is expensive. *Easy Consistency Tuning* (ECT) turns a pretrained diffusion model into a one-step consistency model by fine-tuning, and its smallest setting on CIFAR-10 uses about one million images. We study which distribution of training noise levels (which we call the training path) gives the best one-step generator for a fixed budget, and why. We introduce a novel theoretical framework with the following results. (i) The averaged consistency loss can tend to zero whilst the generator stays wrong. (ii) Adapting a known telescoping bound to the level at which generation starts, we bound the generator's error by quantities that involve only noise levels . (iii) In a simple chain model of training, the least-trained pairs of neighbouring noise levels below limit how fast the generator improves. (iv) A uniform draw of the noise levels gives every such pair nearly the same weight in the averaged loss. From this framework we derive the resistance of a training path, which is computed before any training. We apply these theoretical results to propose *Uniform Consistency Tuning* (**UCT**). Experiments on CIFAR-10 and ImageNet-64 show that **UCT** improves the one-step generation of ECT at small training budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.