acceptodds
Under review as a conference paper at ICLR 2027

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

Abstract

Omnimodal unified models aim to understand text, images, video, and speech and generate coordinated text, image, and speech outputs. The challenge is to extend a single backbone to new modalities efficiently while preserving established capabilities and enabling coordinated cross-modal generation. We present Dynin-Omni, an open-source 8B-scale native omnimodal masked-diffusion model that unifies these capabilities within a single token-level backbone under a shared masked-token prediction objective. We propose a three-stage training recipe for effective omnimodal adaptation that first aligns new modalities without replaying the full task mixture, then applies modality-disentangled merging to recover inherited capabilities. Scheduled padding supervision delays termination learning until semantic grounding is established, enabling flexible-length generation. Moreover, Dynin-Omni is the first masked-diffusion model to natively generate interleaved text, image, and speech within a single reverse diffusion process, enabling mutual conditioning and joint refinement without modality-wise orchestration. Across 20 benchmark settings covering text, image, video, speech, and joint omnimodal generation, Dynin-Omni outperforms prior unified systems on a broad range of directly comparable benchmarks while remaining competitive with modality-specific experts. Stage-wise and controlled analyses show capability recovery and improved length control, while native joint decoding approaches the quality of the best evaluated two-stage generation configuration with a lower sampling budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.