acceptodds
Under review as a conference paper at ICLR 2027

DUET: Dual Uncertainty Evidence Tracking with Visual Tokens and Action Traces in VLA

Abstract

A frozen vision-language-action (VLA) policy may encounter unfamiliar visual content or multiple valid action continuations during deployment. We associate these cases with different uncertainty effects: coverage-related epistemic effects for unfamiliar visual content and action-related aleatoric effects for multiple valid continuations. Existing monitors often merge both into one score and obscure the source of a warning. DUET tracks the two evidence sources separately: a visual evidence branch measures novelty relative to successful demonstrations, while an action evidence branch measures local conditional dispersion at a fixed observation and noisy flow state. Separate references calibrate the two branches into EXECUTE, CAUTION, and HELP, and the visual branch also provides spatial evidence. The visual evidence branch detects unfamiliar observations with frame AUROC and identifies the corresponding regions. The action evidence branch detects local action ambiguity with AUROC . On a physical robot, HELP triggers in unfamiliar-object trials and familiar-object trials, with every unfamiliar alert preceding irreversible motion. DUET shows how a frozen VLA can track the provenance of uncertainty-related evidence and convert complementary evidence into calibrated deployment responses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.