acceptodds
Under review as a conference paper at ICLR 2027

SPIN: Scale-aware Position-unbound Identity Injection for Consistent Visual Autoregressive Generation

Abstract

Visual autoregressive (VAR) models generate images from coarse to fine over a schedule of scales, yet existing methods for subject-consistent generation improve identity consistency across prompts at the expense of pose and layout diversity. Methods that expose source activations to the target also expose their positional information. In VAR models, this locks the target's composition at the coarse scales where layout is determined, tying gains in identity consistency to losses in diversity. We present SPIN, a training-free method that relaxes this trade-off. Its premise is that the source-target state gap is a single object that a transformer block admits through exactly three interfaces, each capable of carrying identity while shedding positional information. State Read Transfer moves this gap into what a layer reads. RoPE-Rewound Subject Attention allows the target to attend to source keys whose rotary phase is rewound on the band that carries position. Pre-Render Write Transfer transfers the per-layer increment, in which everything accumulated before the layer cancels out. Each transfer is placed according to scale- and layer-dependent positional effects. On a 191-group benchmark, SPIN reduces DreamSim on Infinity-2B from 0.235 to 0.182 while retaining 89% of the backbone's pose diversity. It lies above the identity-diversity trend of twelve prior methods, and the same placement recipe improves identity across five tested next-scale backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.