acceptodds
Under review as a conference paper at ICLR 2027

Qwen-Image-Omni: Complex Visual Programs Don't Need External Reasoner

Abstract

Unified multimodal models (UMMs) support multi-turn understanding and generation of interleaved images and text. Complex visual programming requires finer control at each turn: choosing an operation, identifying its target, and recalling historical images with distinct roles while preserving earlier decisions. We present Qwen-Image-Omni, which connects a pretrained Vision-Language Model (VLM) and Diffusion Transformer (DiT) through structured skill calls and a shared visual history. The VLM uses its internal reasoning to select the next operation, assign reference roles, and expand the goal into an executable visual instruction. The same VLM supplies the DiT's conditioning, and each generated image becomes available for subsequent reasoning and execution. We construct histories that preserve dependencies between single-turn skills, then adapt visual execution and prompt expansion in two stages. Autoregressive supervision targets visual prompt content while system instructions specify call syntax. Under matched training and call budgets, the combined skill-and-reference interface improves complex visual design and historical editing over general instructions with natural-language references. Public benchmarks further demonstrate strong interleaved generation, reasoning-informed editing, and general generation and editing capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.