acceptodds
Under review as a conference paper at ICLR 2027

Fusion-to-Branch Distillation for Label-Efficient Multimodal Activity Recognition

Abstract

Label-efficient multimodal learning uses complementary views to reduce the need for labeled data. In multimodal human activity recognition (HAR), this requires a strong complete-view predictor and modality-specific predictors that remain accurate when only one sensor stream is available. Existing semi-supervised methods either create targets from one view or exchange predictions between views, so no branch learns from targets based on both modalities. We introduce fusion-to-branch distillation (F2B), which transfers the complete-view posterior to each modality branch on paired unlabeled examples. A single training run produces one fusion predictor and two single-modality predictors. F2B requires neither a pretrained teacher nor an additional network at inference. We evaluate F2B on five HAR datasets using four label budgets. Compared with supervised training, F2B improves fusion performance in 17 of 20 dataset-budget settings. It improves the direct branches in 38 of 40 evaluations. In 17 of 20 settings, the average branch improvement is also larger than the fusion improvement. Ablations on three datasets show that each branch benefits more from predictions based on the other modality than from its own, supporting cross-modal transfer, although the small number of independent evaluation units limits statistical power. At inference, using one branch requires 44-56% fewer floating-point operations than full fusion. Overall, F2B improves single-modality prediction while retaining the benefits of multimodal fusion in one model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.