acceptodds
Under review as a conference paper at ICLR 2027

OmniDIS: Dichotomous Image Segmentation with Detail-Reference Tokens

Abstract

High-fidelity dichotomous image segmentation requires both selecting the intended foreground and preserving its fine structure. Latent generative models offer strong visual priors, but compressed image representations can weaken boundary evidence, while a caption may match multiple objects. We present OmniDIS, which conditions image-to-mask flow on two complementary sources of evidence. Language–Point Dual-Prompting (LPDP) places semantic descriptions and labeled spatial cues in a shared context sequence, enabling text, point, and joint prompting with one model. Detail-Reference Token Injection (DRTI) converts multi-scale image residuals into learned reference tokens that condition the velocity field before decoding. This design lets target selection and detail cues interact during mask formation, without a separate refinement head. A frozen decoder transmits mask-supervision gradients to both modules, while low-rank adaptation updates the pretrained transformer. On DIS5K, OmniDIS-9B reaches and on DIS-VD with four Euler steps. It surpasses all recent state-of-the-art methods, including FlowDIS and LawDIS, while OmniDIS-4B remains competitive with a substantially smaller generative backbone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.