DreamAR: Knowledge Distillation Sampling Using Next Scale Autoregressive Models
Abstract
We introduce DreamAR, an optimization framework that turns next-scale autoregressive (AR) models into priors for 2D and 3D representations, serving as the discrete AR counterpart to Score Distillation Sampling (SDS). While next-scale AR models rival diffusion in image fidelity and scale favorably, they cannot directly drive SDS, which relies on noise perturbation. Instead, these models operate across discrete multi-scale token maps, predicting categorical distributions over codebook indices in a single forward pass. To leverage this, we propose Knowledge Distillation Sampling (KDS), an objective that minimizes the cross-entropy between the model's predictive distribution and a differentiable tokenization of the rendered asset. Optimization progresses from coarse to fine, advancing to each scale after optimizing the preceding ones, with the scale index serving as the discrete analogue to diffusion timesteps. Because the prior's prediction is fully determined by the rendered view, the prompt and the current scale, KDS supervision involves no sampled noise and avoids the gradient variance that diffusion distillation needs dedicated machinery to reduce. Evaluated on text-to-2D generation (MJHQ-30K) using the Infinity-2B backbone, KDS achieves 24.24 FID, outperforming diffusion distillation baselines (36.55 for the strongest), and even surpasses distillation from SD3.5-Medium despite its lower native generation FID. For both text-to-3D generation and TRELLIS.2-based refinement, KDS performs on par with leading distillation sampling methods CFD and RFDS-Rev across appearance and geometry metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.