Beyond Readability: Rethinking LLM-based MMKGC as Parallel Semantic Ranking
Abstract
Multimodal knowledge graph completion (MMKGC) remains challenging due to structural sparsity, despite its importance for semantic search and reasoning. LLM-based approaches typically address this by formulating MMKGC as a text generation problem. However, existing methods primarily rely on textual inputs and underexploit visual information in multimodal knowledge graphs. Morover, this formulation introduces two key limitations: autoregressive (AR) decoding incurs high latency, and optimizing for readability leads to hallucinations in the entity space. To address the aforementioned issues, we argue that MMKGC is fundamentally a semantic ranking problem within the entity space, rather than a natural language generation task. Driven by this core insight, we propose RIN, a Non-Autoregressive (NAR) parallel inference framework that reframes the LLM-based completion task into a process of semantic ranking over candidate entities. By decoupling the model’s inference trajectory from contextual coherence constraints, RIN not only enables efficient non-autoregressive semantic reasoning but also mechanically circumvents hallucinations typically induced by the pursuit of sequential coherence. Specifically, RIN first employs the Instruction Fine-tuning Module to achieve task reformulation and distribution alignment. Subsequently, the One-shot Generation Module distills core semantic features and filters redundant noise, thereby enabling efficient one-shot parallel inference. Finally, the Hallucination Suppression Module utilizes token-level joint probabilities to perform global scoring of candidate entities. Experiments on three benchmark datasets show that RIN consistently outperforms 22 state-of-the-art methods, achieving both superior accuracy and significantly improved inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.