acceptodds
Under review as a conference paper at ICLR 2027

Which Internal Neural Representations Should Become Spikes? Task-Relevant Spatiotemporal Threshold Adaptation for Audio-Visual Spiking Neural Networks

Abstract

Spiking neural networks (SNNs), regarded as the third generation of neural networks, exhibit rich spatiotemporal dynamics. Audio-visual SNNs (AV-SNNs) can exploit the temporal nature of spike-based computation and event-driven processing for efficient multimodal learning. However, existing AV-SNN studies often leave a more fundamental question implicit: which internal neural representations should become spikes? Different modalities exhibit different task-relevant spatiotemporal structures, making it necessary to preserve highly task-relevant representations while distinguishing redundant ones. We propose Spatio-Temporal Modality-Adaptive Thresholding (ST-MAT), consisting of three components: Task-Relevant Spatiotemporal Support Estimation (TRSSE) estimates task relevance from continuous pre-spike representations and identifies highly task-relevant locations in the spatiotemporal domain. Semantic–Redundancy Balancing (SRB) measures how task-relevant information is distributed over non-redundant spatiotemporal locations and balances semantic retention against redundancy rejection. Slow Bounded Threshold Adaptation (SBTA) learns neuronal thresholds through slow bounded updates to determine which continuous responses should be converted into spikes. Rather than prescribing how many spikes a modality should emit, ST-MAT changes which internal neuronal representations are converted into spikes. ST-MAT achieves 78.47% on AVE, 79.44% on CREMA-D, and 99.54% on UrbanSound8K-AV, improving ST-MAT with fixed threshold by 5.2%, 3.23%, and 1.14%, respectively. It also surpasses the same-architecture ANN on AVE and CREMA-D. Visualizations and ablation experiments show that ST-MAT converts high-task-value internal neuronal representations into spikes while reducing redundancy, thereby substantially improving AV-SNN performance. ST-MAT introduces no additional inference overhead, making it well suited to brain-inspired edge multimodal fusion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.