acceptodds
Under review as a conference paper at ICLR 2027

X-Chat: Any Perception via Multi-Modal Dialogue

Abstract

Multimodal large language models (MLLMs) take a perception request as a turn of dialogue, whether it names a category, describes an object, asks a question, or points at a region. Their outputs are not unified, however: segmentation MLLMs return masks alone, and MLLMs that report boxes write them as text or decode them apart from masks, each covering part of the task spectrum. We present X-Chat, an MLLM that takes multimodal dialogue as its perception interface and covers the segmentation and detection forms of seven task families (generic, open-vocabulary, referring, reasoning, grounded conversation generation, interactive, and visual grounded), fourteen tasks in all. Its language model turns the instruction into condition embeddings that drive one perceptor decoder. Each object query of that decoder regresses a box in parallel with its mask, and a geometric consistency loss couples box and mask. A detection task therefore differs from its segmentation counterpart only in which prediction is returned, and box-only data enter joint training directly. Because the existing interactive and visual grounded benchmarks score masks only, we build and release LVIS-VGD and LVIS-INT for the detection form of these two families. With one set of weights, X-Chat reports results on all fourteen tasks and is compared task by task with X-SAM and with published specialist models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.