acceptodds
Under review as a conference paper at ICLR 2027

DepthFlow-VLA: Routed Low-Noise Flow Refinement for Compact VLA Models

Abstract

Compact vision-language-action (VLA) models commonly allocate the same action-decoding budget to every chunk despite substantial variation in prediction difficulty. We present DepthFlow-VLA, a coarse-to-refine flow-matching decoder in which Stage A first produces a coarse action and a lightweight router selectively invokes restart-style low-noise Stage B. Training combines difficulty-weighted velocity supervision, routing supervision, and a Stage-B objective constructed from replayed Stage-A predictions. On LIBERO, Stage-A-only, adaptive, and uniform-refinement modes achieve 95.9%, 98.3%, and 98.5%, respectively, while adaptive routing uses 7.54 rather than 13.0 decoder evaluations per chunk. On SimplerEnv-WidowX, adaptive and uniform refinement achieve 80.2% and 81.3%, and equal-budget random routing reaches 76.0%. A separately trained vanilla flow-matching control reaches 94.8% on LIBERO and 72.2% on WidowX at six decoder evaluations. On a single H100 with batch size 1, adaptive routing reduces mean end-to-end latency from 59.2 to 42.6 ms relative to uniform refinement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.