acceptodds
Under review as a conference paper at ICLR 2027

VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers

Abstract

Visual in-context learning (V-ICL) uses a source–target image pair to specify a transformation for a new query. A central challenge is to support structured prediction, restoration, and open-domain editing through this interface, despite their different output structures and appearances. We present VIRAL, a V-ICL framework for pre-trained Diffusion Transformers (DiTs), instantiated with Qwen-Image-Edit. VIRAL encodes the exemplar and query as separate image sequences, distinguishes their roles through 3D-MSRoPE, and uses MoE-LoRA to accommodate heterogeneous transformations while keeping pre-trained weights frozen. We construct a dataset of transformation-consistent exemplar–query quadruplets through standard-task pairing and analogous edit mining and synthesis. The resulting model follows visual demonstrations without task-specific text instructions or test-time parameter updates. Across eight task groups, VIRAL outperforms four V-ICL baselines on all 14 reported metrics. It also performs open-domain editing and generalizes to held-out styles and lineart generation, which are excluded from VIRAL training. These results show that one pre-trained DiT can support a shared visual-demonstration interface across diverse tasks and edits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.