When Foresight Drifts: Progress Aligned Resynchronization for Cloud–Edge Robotic Control
Abstract
Deploying billion-parameter vision-language-action (VLA) policies on mobile robots creates a systems tension: semantic reasoning benefits from cloud GPUs, whereas closed-loop control requires responsive local execution. We identify two sources of foresight drift in cloud–edge robotic control. First, uniform temporal prediction allocates representational capacity by elapsed duration rather than semantic significance, creating a pace–progress mismatch that underrepresents critical procedural transitions. Second, physical changes can invalidate predicted continuations even when the semantic stage remains relevant, creating a state-foresight mismatch. We propose Foresight Resync (ForRes), which integrates semantic-target-conditioned cloud prediction, progress-based edge cache addressing, and value-guided asynchronous refresh. In the cloud, a JEPA-based predictor forecasts the next semantic milestone from causal execution history, and a lightweight adapter conditions the world action model on this target while preserving its native prediction budget. On the edge, recurrent addressing matches observed execution to cached semantic anchors, retrieving latent guidance according to procedural progress rather than elapsed time. Target-support gating suppresses incompatible guidance, while an alignment-conditioned critic requests refresh when its expected benefit exceeds querying costs; invalid or exhausted caches trigger refresh directly. Edge policy maintains observation-only closed-loop control while awaiting cloud responses. Evaluated with two cloud backbones and an X-VLA edge policy, ForRes achieves up to 78.73% average success on RoboTwin 2.0, improving over Latent-to-Action by 4.20 percentage points with 31.11% fewer cloud calls, and reaches 76.00% success on real-world dual-arm tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.