H-MPO: Hierarchical Multi-Response Preference Optimization with Modality-Guided Decoding for Omni Large Language Models
Abstract
Omni-modal large language models (OLLMs) exhibit capabilities in processing and integrating audio-visual modalities but remain vulnerable to cross-modal hallucinations. To mitigate this, existing methods employ binary preference learning by contrasting relevant responses against completely disconnected ones. This coarse-grained method overlooks intermediate hallucination states, leading to sparse supervision that struggles to address complex cross-modal conflicts. In this paper, we introduce Hierarchical Multi-Response Preference Optimization (H-MPO), which shifts preference learning from a binary comparison to a multi-hierarchy process. To obtain hierarchical preference data, we design an annotation-free Modality-Guided Decoding pipeline that synthesizes a four-level preference ranking set, capturing distinct degrees of modality faithfulness. To leverage the constructed data structure, H-MPO incorporates asymmetric margin constraints to jointly optimize the preference chain. We further propose a token-level reweighting mechanism to strengthen the supervisory signal for modality-sensitive tokens. Extensive experiments on hallucination benchmarks and general tasks demonstrate that H-MPO mitigates cross-modal hallucinations while preserving the general perception and reasoning capabilities of models. Our code and dataset will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.