acceptodds
Under review as a conference paper at ICLR 2027

MetricFly: Numerical Instruction Following via Action-Conditioned Calibration for UAV Vision-Language-Action Models

Abstract

Aerial vision-language-action (VLA) policies can generate task-relevant UAV motion, but often fail to meet numerical requirements such as a specified orbit radius or flight altitude. This paper studies how to improve numerical execution by exposing the current metric error to the policy and calibrating the motion it actually generates. We introduce MetricFly, a two-stage framework that first conditions a policy on the target, current task metric, and signed residual, then uses a lightweight external calibrator to apply bounded geometric corrections to the sampled action chunk. The calibrator is trained with supervision from an analytic geometric teacher while the fine-tuned policy remains frozen during calibrator training. We evaluate MetricFly on HUGE-Bench-derived Orbit-R and Orbit-H tasks with 3DGS-rendered visual feedback and externally supplied reference geometry. Explicit numerical conditioning reduces trajectory task-metric MAE by and , respectively, compared with text-only conditioning. Adding action-conditioned calibration raises full-orbit completion from to over the metric-conditioned policy alone. Replacing the calibrator's input with a mismatched action sample reduces calibration precision, showing the value of conditioning corrections on the motion being executed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.