acceptodds
Under review as a conference paper at ICLR 2027

DINOv3 Perceptual Loss with Token and Correlation Matching for Spike Image Reconstruction

Abstract

Spike image reconstruction aims to recover visually meaningful images from asynchronous spike streams, but remains challenging in neuromorphic vision. Pixel-wise losses often fail to capture perceptual similarity, resulting in overly smooth or unnatural images. Perceptual supervision is therefore particularly important for this task, as it provides the structural, semantic, and texture information that cannot be directly inferred from raw spike signals. We propose a DINOv3 perceptual loss specifically designed for spike image reconstruction through complementary token and correlation matching. Built on self-supervised DINOv3 representations, class-token matching preserves global semantic consistency, patch-token matching retains fine-grained local details, and correlation matching maintains spatial structure and texture coherence. Together, these components provide comprehensive perceptual supervision at both the token and correlation levels. The proposed loss is used only during training and introduces no additional inference cost. Extensive experiments demonstrate that the proposed DINOv3 perceptual loss consistently improves CNN-, Mamba-, and Transformer-based reconstruction models and achieves state-of-the-art performance on the Spike-REDS and Spike-X4K benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.