acceptodds
Under review as a conference paper at ICLR 2027

ScanAR: Parallel Row-Scan Decoding for Efficient Image Generation in Unified Autoregressive Models

Abstract

Unified multimodal autoregressive (Unified AR) models represent text and images in a shared token sequence and generate all modalities with a common next-token prediction objective. However, this raster-order decoding requires \(HW\) sequential decisions for an \(H \times W\) image, overlooking the inherent two-dimensional structure of visual tokens and creating a major inference bottleneck. To address this problem, we propose ScanAR, a row-scan parallel autoregressive method for efficient image generation within Unified AR models. Instead of modifying the generation head or introducing additional parameters, ScanAR accelerates decoding by redefining the predecessors of visual tokens, enabling seamless integration with existing Unified AR models. Specifically, the first visual row follows the standard left-to-right dependency, while each subsequent row is generated in parallel by conditioning each token on its vertically aligned predecessor in the previous rows. For nonvisual tokens, the standard autoregressive rule is unchanged. In this way, ScanAR reduces the sequential decoding from \(HW\) to \(H+W\), supporting both training from scratch and lightweight post-training without modifying the backbone, vocabulary, or output head. Extensive experiments on Emu3.5 demonstrate that ScanAR improves the GenEval score by 1.88 points and achieves a \(27.7\times\) speedup for \(512\times512\) image generation through post-training. On U0, ScanAR accelerates image generation while retaining its interleaved and video generation capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.