acceptodds
Under review as a conference paper at ICLR 2027

Ontology-Harnessing Multi-Agent System for Abstractive Visual Understanding

Abstract

Multi-modal large language models (MLLMs) have made rapid progress in visual understanding, yet abstractive visual representations remain challenging because their interpretation depends on recovering entities, relations, and structural constraints beyond visual appearance alone. In this paper, we propose the Ontology-Harnessing Multi-Agent system (OHM), a training-free framework that decomposes abstractive visual understanding into perception, validation, and task-specific reasoning around an explicit symbolic graph. An ontology provides shared semantic and structural knowledge for intermediate validation and disambiguation, while gradually incorporating reliable observations during inference to support subsequent reasoning. Extensive experiments across multiple tasks and model backbones demonstrate consistent improvements over direct inference, while further revealing the complementary roles of reliable structure recovery and effective downstream reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.