acceptodds
Under review as a conference paper at ICLR 2027

-EGO: Proxies as mediators for Poly-modal-Induced Egocentric Video Representation Learning

Abstract

Egocentric RGB video provides an inherently limited view of human actions, often obscuring body motion and surrounding scene context. Privileged information from complementary modalities, exocentric views, and foundation models can provide richer supervision during training, but is often unavailable or impractical at inference. To this end, we introduce -EGO, a poly-modal-induced egocentric RGB encoder learned through PROXYDIST, a multi-teacher distillation framework that uses proxies as mediators to consolidate knowledge from nine heterogeneous teachers spanning viewpoints, modalities, and foundation model representations. PROXYDIST learns representation-specific proxies that translate heterogeneous knowledge into a homogeneous egocentric space rather than directly distilling from incompatible teacher feature spaces. Then, the proxies are merged through a learned convex initialization, followed by Selective Proxy Distillation (SPD), which selectively distills from the most reliable proxies for each training sample. The resultant -EGO achieves state-of-the-art performance across action recognition, video retrieval, and action segmentation on three challenging ego-exo benchmarks, outperforming direct multi-teacher distillation while requiring only egocentric RGB video and no additional computational overhead at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.