acceptodds
Under review as a conference paper at ICLR 2027

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Abstract

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we introduce , a stronger multimodal diffusion language baseline that integrates diffusion language models (DLMs) with a visual encoder, achieving state-of-the-art performance among open-source diffusion-based MLLMs. Then, by leveraging the parallel decoding nature of DLMs, we further propose an efficient prompting mechanism that enables simultaneous perception of multiple masked regions, allowing the model to generate region descriptions in parallel at both the sequence and token levels. This design significantly improves inference efficiency compared with existing approaches that process regions sequentially. To better evaluate the parallelism property of visual perception capability for DLMs, we construct a new llel etailed ocalized aptioning Benchmark () by scaling the DLC-Bench to include multiple region masks per image, enabling joint evaluation of both caption quality and inference speed. Experiments demonstrate that PerceptionDLM maintains competitive region captioning performance while achieving substantial speed improvements for multi-region perception tasks. Our results highlight the potential of diffusion-based multimodal language models for efficient parallel visual perception. Code and model will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.