acceptodds
Under review as a conference paper at ICLR 2027

Peregrine: Latent-Future-Based Action Correction for Vision-Language-Action Policies in Contact-Rich Dynamic Manipulation

Abstract

Action chunks amortize the online inference cost of vision-language-action (VLA) policies by generating multistep action sequences, but robot and object motion can render later actions in a chunk stale before execution. We propose Peregrine, a lightweight framework that conditions action correction on latent representations predicted for the target execution time. A Dynamic Tokenizer learns structured historical dynamics through robot-motion and region supervision, encoding observed changes in the robot and objects into compact dynamics tokens. A Future Predictor combines these tokens with the latest observation, the action timeline, and temporal conditions to predict latent representations at the target action’s execution time. A Residual Corrector uses the execution-time prediction and the corresponding base action to generate a bounded residual correction, without requiring futureimage generation or complete action replanning. The pretrained flow-based VLA remains frozen. All auxiliary modules are trained offline, and their weights remain fixed during deployment, while predictions and corrections are updated as new observations arrive.We evaluate Peregrine in dynamic simulations and real-robot manipulation, assessing task success rates and future latent prediction accuracy, and examining component contributions through ablation studies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.