acceptodds
Under review as a conference paper at ICLR 2027

Activate Sparse, Predict Continuously: Brain-inspired Ternary Sparse Activation with Predictive Coding for Multi-object Tracking

Abstract

Multi-object tracking (MOT) must detect objects and preserve their identities in continuous video streams under tight computational budgets. Spiking neural networks (SNNs) promise low-power computation, but their sparsity lives along the temporal axis: they encode information by accumulating events over timesteps, which mismatches the instantaneous spatial discrimination and continuous state estimation that online MOT demands, and occlusion-induced event loss destroys the features needed to re-identify a target. To address these issues, this paper proposes an efficient online tracking framework inspired by the sparse coding and predictive cognition mechanisms of biological vision, termed Activate Sparse, Predict Continuously (ASPC). First, we propose a ternary sparse activation network that gates neural responses dynamically and encodes features with ternary sparse activations (-1/0/+1), shifting sparse computation from the temporal domain to a ternary sparse activation-based feature representation space; this mechanism preserves the directional discriminative information of single-frame features while enabling adaptive feature selection. Furthermore, we design a prediction-driven cognitive association module that formulates the tracking process as dynamic inference based on historical state priors, thereby maintaining identity continuity through internal states when visual information is insufficient. Experiments on MOT17, DanceTrack, and SportsMOT show that ASPC matches or surpasses the strongest full-precision trackers in HOTA and AssA while using a single forward pass: on MOT17 it attains HOTA 59.3 and AssA 58.2 at 25.7 mJ per frame, against 1105 mJ for the strongest full-precision baseline, and it exceeds an identically configured spiking counterpart by 29.3 HOTA points on MOT17 and 22.5 points on DanceTrack. Its activation sparsity adapts to scene density (48-59% zero activations), with per-frame inference energy of only 25.7 mJ, 11 mJ, and 19 mJ on the three datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.