acceptodds
Under review as a conference paper at ICLR 2027

EAR: Efficient AutoRegressive Text-to-Image Synthesis from a Game-Theoretic View

Abstract

Conventional autoregressive text-to-image synthesis approach rely on next-token prediction, which would hinder inference efficiency due to their sequential, predefined token generation order. To enable high parallelization without sacrificing generation quality, we revisit autoregressive text-to-image models from a theoretical perspective by analyzing the generation order of visual tokens. We also propose the Efficient AutoRegressive text-to-image synthesis model (EAR), which provides a principled explanation of how visual tokens are generated in autoregressive models, while still supporting fine-grained image synthesis with efficient inference. Specifically, drawing inspiration from Game theory, we suggest employing Shapley value to evaluate the interactions between regions that arise from different order steps. Moreover, we present a novel text-to-image synthesis framework that facilitates rapid generation, with the guidance of both global representation and intermediate local embedding extracted from a pretrained text encoder. Extensive experiments, including quantitative evaluations and qualitative visual comparisons, on the public benchmarks show that our EAR model achieves lower latency than previous autoregressive text-to-image models without sacrificing image quality. We hope that the EAR model will offer novel insights into the interpretability of autoregressive models, thereby providing valuable benefits to researchers in this field.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.