acceptodds
Under review as a conference paper at ICLR 2027

MAPO-Net: MLLM-guided Artifact Perception and Feedback Optimization Network for HDR Reconstruction

Abstract

Despite remarkable advances in dynamic High Dynamic Range (HDR) reconstruction, artifacts such as underexposure, overexposure, ghosting, and motion blur still persist severely when reconstructing HDR images for dynamic scenes. Existing methods mainly rely on implicit feature alignment and lack explicit semantic understanding, making it difficult to robustly identify and restore complex local image degradation. To address this limitation, we propose MAPO-Net, a novel framework that bridges high-level semantic reasoning and low-level image reconstruction. For the first time, we introduce the high-level reasoning capability of Multimodal Large Language Models (MLLMs) into the feedback loop of HDR reconstruction. Our core idea is to convert artifact restoration into a "text-guided detection and penalization" pipeline. Specifically, MAPO-Net consists of two critical stages: Extensive experiments demonstrate that our method significantly improves artifact suppression and detail preservation under dynamic scenes, achieving superior performance on real-world HDR benchmarks. Our framework highlights the effectiveness of integrating MLLM-based reasoning with low-level vision reconstruction tasks. 1.Cross-Modal Artifact Perception: We freeze the MLLM (Qwen3-VL) to act as a diagnostic module that outputs semantic tokens describing various artifacts. A lightweight cross-modal detector then aligns these textual priors with visual features to accurately segment artifact regions. 2.Perception-Driven Feedback Optimization: The predicted artifact masks are converted into adaptive weight maps. Via artifact-aware loss, the base HDR network is compelled to focus on optimizing these defective regions flagged by the MLLM during training. Furthermore, existing public datasets suffer from limited diversity of challenging samples for the aforementioned artifacts. To mitigate this issue, we contribute 2,300 real-world LDR scenes and construct a brand-new HDR dataset comprising these 2,300 hard samples. Extensive experiments conducted on our newly built dataset and existing public benchmarks demonstrate that the proposed MAPO-Net can effectively suppress artifacts that cannot be eliminated by prior approaches, achieving substantial improvements over state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.