acceptodds
Under review as a conference paper at ICLR 2027

ARGO-VLA: Calibrated Gating and Escalation for Correcting and Orchestrating Frozen VLA Policies in the Laboratory

Abstract

Vision-language-action (VLA) policies fine-tuned for laboratory skills fail at two timescales. Within a skill they miss millimeter tolerances: on AutoBio (Lan et al., 2026), a fine-tuned succeeds on at most 43% of episodes of any medium task and 15% of any success-scored hard task. Across a protocol they cannot order steps, withhold hazardous operations, or notice that a step failed. We propose ARGO-VLA, built around one calibrated novelty score used at both timescales. At the control rate the score masks a bounded residual, trained by behavior cloning and PPO on a chunk-level residual Markov decision process and restricted to its training support by a k-nearest-neighbor trust gate, that corrects each action chunk of a frozen . At the protocol-step rate the same score, thresholded by split conformal prediction, halts a skill and escalates to a typed state machine of LLM agents that plans, validates and risk-tags, requests human approval, executes and reviews protocol steps; the threshold bounds the false-escalation rate per nominal skill call. We introduce AnomalyBio, which extends AutoBio with multi-step protocols, injected failures and hazardous or compositionally infeasible goal probes judged by simulator ground truth, and ablate the trust gate, anomaly mask, escalation, validator critique and LLM reviewer. On transfer, the residual reaches 48.7% success versus 40.7% for AutoBio's published , a comparison across different checkpoints and episodes, not a controlled one. Escalation detects 52% of monitor-reachable injected failures at 1.8% false escalations per call, and moves protocol success under injected failures from 45.3% to 48.0%, a gain within one standard error.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.