acceptodds
Under review as a conference paper at ICLR 2027

Cross-Scale Aligned Supervision for Training GANs

Abstract

Modern GANs often use multi-scale adversarial supervision to provide direct realism feedback to intermediate generator outputs. However, we identify an important limitation of standard scale-wise supervision: each intermediate output is independently optimized for realism at its own resolution, without explicitly coordinating outputs across generator stages. As a result, realistic intermediate outputs need not remain aligned with the final output from the same generated sample. This motivates treating per-scale realism and cross-scale consistency as complementary objectives for multi-scale adversarial generation. Based on this observation, we propose CAT, a Cross-scale Aligned Transformer, which preserves scale-wise adversarial supervision for direct per-scale realism feedback while introducing a simple generator-side consistency regularization to align intermediate outputs with the final output. We further show that jointly exposing multiple scales to the discriminator improves cross-scale consistency but substantially degrades generation quality, supporting our design of enforcing cross-scale consistency separately on the generator side. In a matched 40-epoch CAT-H/2 comparison on ImageNet-256, the proposed consistency regularization improves FID-50K from 4.35 to 1.76. CAT-H/2 achieves an FID-50K of 1.56 with a single generator evaluation after only 60 training epochs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.