acceptodds
Under review as a conference paper at ICLR 2027

KI-VLA: Kinematics-Informed Action Decoding for Flow-Matching Robot Policies

Abstract

Flow-matching robot policies can generate operational-space action chunks without explicitly modeling the underlying joint-space motion required for execution. As a result, decoded chunks can contain jitter, micro-reversals, and locally inconsistent commands that degrade closed-loop performance. We introduce KI-VLA (Kinematics-Informed VLA), a predictor-corrector decoder that treats the flow policy's intermediate end-effector action chunks as noisy kinematic observations and uses them to infer a smooth latent joint-increment trajectory. A kinematic observation model and Kalman/RTS smoothing infer temporally coherent joint motion, which is mapped back to action space via bounded trust-region updates. By interleaving these corrections with flow-matching steps, KI-VLA refines intermediate proposals while keeping the base policy frozen. On RoboTwin 2.0, KI-VLA improves the baseline from to in the benchmark's Easy setting and from to in its Hard setting. On real hardware, clean-condition success is comparable ( to ), while success under a controlled disturbance injected into the action proposal rises from to . Further studies show that the gains are concentrated in the few-step decoding regime and can diminish under severe measurement-covariance or process-noise mis-scaling, clarifying the operating regime in which decoding-time kinematic correction is most useful.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.