FReD-VLA: A Vision-Language-Action Model with Flow Matching-Initialized Diffusion
Abstract
Deploying vision-language-action (VLA) models for bimanual manipulation in real-world environments requires robots to generate accurate actions with low latency to respond to changing surroundings. State-of-the-art methods usually adopt diffusion- or flow matching-based approaches for continuous action generation. Despite promising results, balancing accuracy and latency still remains challenging. Diffusion models can generate accurate actions, but repeated denoising increases computation and latency. One-step flow matching is efficient, but its actions tend to be inaccurate for fine-grained manipulation. In this paper, we propose a Vision-language-Action model with Flow Matching initialized Diffusion (FReD-VLA), which combines fast flow-based generation with diffusion-based correction. Given visual observations, language instructions, and robot states, a flow matching model first generates a coarse action sequence. The coarse action sequence is then refined by a diffusion model through Action Distribution Calibrator, enabling refinement from a well-initialized intermediate diffusion state rather than the fully-corrupted beginning state. Empirically, on Maniskill PickCube-v1, our model achieves a mean minimum end-effector-to-cube distance of 10.8 mm, with reduction than the 43.7 mm measured for pure Diffusion. In offline evaluation, one diffusion step reduces the joint mean absolute error of a one-step flow proposal from to . This is a 47.3% reduction. With one flow evaluation and one diffusion evaluation, FReD-VLA achieves an action-generation latency of 37.0 ms, excluding shared multimodal encoding. This provides a speedup over five-step diffusion with comparable action accuracy. Real-robot tests further demonstrate dynamic bimanual assembly after adaptation with about 20–30 human demonstrations per behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.