acceptodds
Under review as a conference paper at ICLR 2027

GeoMFF: Reconstructing a Latent Sharp 3D Scene for Misaligned Multi-Focus Image Fusion

Abstract

Multi-focus image fusion (MFIF) aims to integrate focused information from different focal planes into an all-in-focus image. However, existing methods typically perform focus estimation and fusion in the 2D image space, assuming pixel-wise correspondence among source images. They therefore struggle to handle spatial misalignment caused by camera motion and depth-dependent parallax during practical acquisition. To address this issue, we are the first to formulate misaligned MFIF as the recovery of a latent sharp 3D scene from multiple defocused observations, and propose GeoMFF, a geometry-aware fusion framework. GeoMFF initializes a 3D Gaussian scene using camera parameters and point clouds predicted by VGGT, and jointly optimizes the scene representation, camera poses, and view-dependent spatially varying defocus model to aggregate focused information from different focal planes in 3D space. Within this framework, the Gaussian-Initialized Mixture Blur Network (GMBN) compactly models spatially varying defocus by adaptively combining a shared set of learnable blur kernels. We further construct a real-world misaligned multi-focus image dataset comprising two-image pairs and seven-image sequences. Experiments on our dataset and multiple public benchmarks demonstrate that our method achieves state-of-the-art performance on misaligned image pairs, remains highly competitive on aligned benchmarks, and extends effectively to multi-focus image sequences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.