CoFiTok: Coarse-to-Fine Denoising Tokens for Pixel-Space Diffusion
Abstract
Diffusion networks predict an image-sized noise field at each evaluation. We in- troduce CoFiTok, which organizes this field as an ordered sequence of continuous token contributions. Each token is synthesized independently through a shallow bias-free map, giving a zero token an exact zero-component meaning. Accumu- lating the components yields partial clean-image estimates at a fixed noisy state and diffusion time. Prefix and component losses train these estimates along a pre- scribed coarse-to-fine denoising path. We relate pixel-space prefix errors to noise- space errors and use fixed-component permutations to evaluate ordering while preserving the full prediction. Across eight dataset settings, CoFiTok consistently improves path alignment. Two-seed studies with 20,000 training steps reduce path AUC from 1.0098 to 0.0378 on Tiny ImageNet and from 0.9984 to 0.0379 on ImageNet-64 HF, with endpoint-MSE increases of 3.4% and 5.0%. On ImageNet- 256, the learned four-component order ranks first among all 24 permutations for both seeds. Ablations identify complementary effects of prefix and component su- pervision, alongside mixed endpoint and generation results. CoFiTok provides a structured interface for studying partial denoising within a diffusion prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.