NextGAN: Towards Simple and Scalable GAN Training for Image and Video Generation
Abstract
Generative adversarial networks (GANs) enable one-pass sampling, but their critics are trained to separate real and generated samples rather than to provide a useful correction field. We present NextGAN, which directly supervises the critic's input gradient. Independent real–generated pairs are interpolated in frozen feature spaces, and the gradient is regressed onto their displacement. Random-pair regression aggregates into a distribution-dependent field that vanishes when the conditional feature distributions match; for a fixed generator, its population objective has an exact adversarial dual with a gradient-energy penalty. Without auxiliary losses, online teacher scores, or trajectories, the same objective trains pixel- and latent-space image generators and Wan2.1 video generators up to 14B parameters, remaining stable across one- and few-step sampling budgets. On ImageNet 256*256, NextGAN achieves FID=2.87 with pixel-space JiT-B and FID=1.54 with latent-space DiT-XL/1 at one evaluation, while the adversarial form reaches FID=1.36. On text-to-video generation, the 1.3B model achieves the highest VBench total score among the compared 1.3B methods at one- and two-step budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.