acceptodds
Under review as a conference paper at ICLR 2027

UniOwl: Bridging Multimodal Understanding and Generation with Adaptive Conditioning

Abstract

Unified multimodal models aim to support visual understanding and image generation within a single architecture, but diffusion-based generation introduces a key trade-off: while semantic guidance should adapt across denoising steps, invoking a large language model at each step is computationally prohibitive. Existing methods rely on either static conditions or costly step-wise language computation. We introduce UniOwl, a unified multimodal framework for joint understanding and generation with Adaptive Dual-Channel Conditioning. UniOwl first extracts step-invariant semantic condition tokens from LLM hidden states, computed once and shared across all denoising steps, then refines them at each step with lightweight latent- and timestep-aware residuals. In parallel, visual condition tokens are derived from the current denoising latent via a shared visual encoder. The fused semantic and visual tokens provide dynamic conditioning signals for the diffusion model, while generation supervision on the visual pathway ensures consistency between generation and understanding. To further encourage generation-understanding consistency, we introduce a cycle consistency objective that feeds generated images back into the understanding pathway and encourages their output distributions to be consistent with the original predictions. Extensive experiments show that UniOwl achieves superior generation quality and semantic consistency over existing unified models, while maintaining competitive multimodal understanding performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.