LightMGT: A Compact Masked Generative Transformer for Visual Generation
Abstract
Masked generative transformers (MGTs) synthesize images by predicting discrete visual tokens in parallel, but existing models offer limited efficiency advantages over diffusion models. We therefore propose LightMGT, a compact masked generative transformer for image generation and editing. Its 148M backbone reduces model size through a compact double-stream design and factorized token prediction. We introduce MGT-OPD to combine generation and editing capabilities in one model by matching specialist teachers’ token predictions on student-generated masked states. Antithetic Posterior Matching (APM) then distills iterative decoding into one step using a pair of masks and an updated auxiliary model, yielding LightMGT-Turbo. LightMGT achieves 0.70 on GenEval, 78.32 on DPG-Bench, and 6.49 on GEditBench-EN. On GenEval, it scores 0.70 versus 0.67 for the 12B-parameter FLUX.1-dev, while delivering 8.8× higher throughput. With a single decoding step, taking about 0.235 s, LightMGT-Turbo achieves scores of 0.58 on GenEval, 71.30 on DPG-Bench, and 6.21 on GEditBench-EN. MGT-OPD combines text-to-image generation and editing capabilities. In the LightMGT experiments, it improves GenEval and GEditBench-EN by 16.7% and 4.5% relative to weight merging, and by 11.1% and 7.1% relative to joint training, respectively. APM also improves one-step distillation. On native MaskGIT ImageNet-256, it reduces FID by 1.2% and increases IS by 4.8% relative to Di[M]O. Together, this work contributes to the community a compact backbone, capability composition, and one-step distillation for image generation and editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.