EgoExo-ViewCL: View-Labeled Continual Action Learning with Stable Analytic Anchors and View-Factorized Adaptation
Abstract
Continual video learning is usually evaluated as a sequence of new action classes, yet an embodied learner also moves between first- and third-person views. A model may preserve action semantics within an observed view yet lose cross-view transfer, or adapt by overwriting useful pretrained structure. We introduce EgoExo-ViewCL, a view-labeled continual action benchmark built from EgoMe, Assembly101, and Charades-Ego. Its three protocols isolate view shift under a fixed label space, joint class-and-view expansion, and sequential paired-view arrival. The benchmark preserves native single- and multi-label targets and evaluates both final recognition and continual dynamics, including view-conditioned interference, view-order robustness, and paired-view consistency. We further propose VFCT, which combines an order-stable analytic anchor, bounded semantic adaptation, and cumulative class–view prototype score mixing. Controlled evaluation keeps the observation set fixed when measuring view order, averages five class orders, and crosses paired-view arrival direction with true and pseudo pairing. Across the three datasets, VFCT consistently improves recognition under view switches, joint class–view expansion, and paired-view transitions. Continual and protocol-specific diagnostics show that retaining shared action evidence and adapting view-dependent cues are complementary throughout the stream. These results establish viewpoint order, pairing, and their interaction with class arrival as important axes for continual action learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.