EV-JEPA: Elastic V-JEPA for Any Device and Cross-Device Inference
Abstract
Foundation models like V-JEPA are increasingly deployed across a heterogeneous edge–cloud continuum spanning smart glasses, phones, robots, and servers. Yet, they are typically trained as rigid, fixed-capacity networks bound to one device, supporting neither reconfigurable compute budgets nor collaborative execution. Existing approaches address only one axis at a time: slimmable and NAS-based methods provide width elasticity but assume single-device inference, whereas split-computing methods enable cross-device collaboration only for fixed partitions and static capacities. We introduce Elastic V-JEPA (EV-JEPA), a single model trained once and deployable everywhere along the continuum. Its width is reconfigurable, and it is natively partitionable across heterogeneous devices. At its core is Transition-Aware Triplet Training (T3): at every step, a pair of widths is trained both standalone and as endpoints of a hybrid configuration that transitions between them in a split layer, so that the same weights serve the execution of single-width and cross-width. Across various benchmarks, EV-JEPA falls within 1% of independently trained baselines while improving over previous elastic training approaches by up to 11%. Cross-device hybrid execution recovers to within 2% of the highest-capacity standalone model and supports many-to-one, any-to-any, and variable-split configurations. Finally, we show that the elastic principle extends beyond the backbone to an elastic token compression mechanism that exposes controllable trade-offs among on-device computation, latency, and accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.