-EGO: Proxies as mediators for Poly-modal-Induced Egocentric Video Representation Learning
Abstract
Egocentric RGB video provides an inherently limited view of human actions, often obscuring body motion and surrounding scene context. Privileged information from complementary modalities, exocentric views, and foundation models can provide richer supervision during training, but is often unavailable or impractical at inference. To this end, we introduce -EGO, a poly-modal-induced egocentric RGB encoder learned through PROXYDIST, a multi-teacher distillation framework that uses proxies as mediators to consolidate knowledge from nine heterogeneous teachers spanning viewpoints, modalities, and foundation model representations. PROXYDIST learns representation-specific proxies that translate heterogeneous knowledge into a homogeneous egocentric space rather than directly distilling from incompatible teacher feature spaces. Then, the proxies are merged through a learned convex initialization, followed by Selective Proxy Distillation (SPD), which selectively distills from the most reliable proxies for each training sample. The resultant -EGO achieves state-of-the-art performance across action recognition, video retrieval, and action segmentation on three challenging ego-exo benchmarks, outperforming direct multi-teacher distillation while requiring only egocentric RGB video and no additional computational overhead at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.