Better Positions or Different Values? Rethinking Confidence Ranking in Guided Masked Image Generation
Abstract
Confidence ranking in masked image generation commits the positions whose candidate tokens are most probable, and its advantage over random positions depends on the sampling temperature. Hayakawa et al. (2026) explained it as implicit cooling on an unconditional image model. With assumptions about complete generation, we turn this cooling interpretation into a temperature rule for text-to-image models with classifier-free guidance. We develop the rule on one model and test it on three more. As the rule predicts, confidence ranking helps above the best tested temperature for random positions but brings no significant improvement at that temperature. On MMaDA-8B, confidence ranking improves Fréchet Inception Distance (FID) by up to 51.8 at a high temperature, whereas the released confidence ranking worsens FID by 14.1 at the best tested temperature for random positions. After each sampler receives its own best tested temperature, confidence ranking has no significant FID advantage in three model families at released step counts. We extend the single-step analysis to many simultaneous commits and to position scores computed from probabilities at temperature one. This improves predictions of committed values and provides a starting point for tuning ranking strength. A controlled change to position scores shows that selection can weaken the guidance reflected in committed values even when candidate-value sampling stays fixed. These findings call for temperature-calibrated baselines and guidance controls when the scoring rule changes effective guidance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.