Dynamic Merging Mutually-Distilled Multi-View Teacher Network for RGB-Event based Representation Learning
Abstract
Multi‑modal fusion algorithms combining conventional frame‑based cameras and event cameras have recently attracted substantial research attention. Typically, existing approaches convert event streams into frames, voxels, or tensors, and then learn the required spatiotemporal information from these transformed event representations for fusion with conventional frame‑camera data. Obviously, such paradigms only exploit the limited representational capacity of event data. Moreover, standard training pipelines fail to fully leverage diverse model parameters obtained throughout the learning process. In this paper, we propose a novel mutual distillation method leveraging multi‑view teacher network weights to better integrate multiple event representations and achieve model‑parameter level fusion. Specifically, we employ multiple encoders to extract different types of event representations and perform mutual knowledge distillation at both the feature‑level and logit‑level in this process. Afterwards, we perform dynamic hierarchical model merge strategy on the model parameters from the first stage to yield a unified multi‑modal RGB‑Event representation model. Experimental results across multiple event-based datasets demonstrate that our proposed framework achieves significant performance improvements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.