SparseGRN: Accelerating Generative Refinement Networks with Adaptive Sparse Queries
Abstract
Generative Refinement Networks (GRNs) synthesize images by repeatedly revisiting the complete token map, enabling global error correction but incurring substantial inference cost. Recomputing every visual token at every refinement step introduces redundancy as image content stabilizes. We present SparseGRN, a training-free framework that accelerates pre-trained GRNs through refinement-aware adaptive query routing. Our analysis reveals that semantic structure forms before local detail refinement is complete, bit uncertainty identifies tokens that remain sensitive to refinement, and sparse query updates benefit from complete key/value coverage. Building on these observations, SparseGRN matches global uncertainty to offline dense reference statistics to determine a prompt-dependent routing onset. During routed refinement, a budget-adaptive controller combines prompt-conditioned response estimates with an offline dynamic-programming table to allocate queries across the remaining steps. High-uncertainty tokens receive fresh attention and feed-forward updates, while inactive positions reuse cached block outputs and keys and values are recomputed over the complete context. This design preserves GRN's sampling schedule and full-map sampling procedure while reducing per-step computation. On an NVIDIA A100 GPU, SparseGRN achieves a speedup, reducing latency from 20.96 s to 14.34 s with a GenEval score of about 0.77 and a DPG-Bench decrease of only 0.08 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.