acceptodds
Under review as a conference paper at ICLR 2027

AEGIS: Action-Endpoint Geometry for Nuisance-Invariant and State-Equivariant VLA Models

Abstract

Robust vision-language-action (VLA) models must remain stable under task-equivalent language and task-irrelevant visual changes while adapting their actions to control-relevant changes in robot state, object configuration, or task progress. Per-sample mean squared error (MSE) supervision anchors individual predictions to demonstrations, but its fixed quadratic metric does not adapt residual weighting or explicitly constrain how predictions should relate across these conditions. We present AEGIS, a demonstration-anchored action-endpoint geometry for VLA models that predict action chunks from single-timestep observations. It combines three complementary objectives in a shared action field: Bounded Heteroscedastic Flow Matching (BHFM) uses bounded heteroscedastic scales to adapt residual weighting; Nuisance Endpoint Invariance (NEI) encourages endpoint agreement under action-preserving language and visual variations; and State Endpoint Equivariance (SEE) encourages local endpoint displacements to match demonstrated action differences across progress-matched states. AEGIS achieves 91.58% success on LIBERO-Plus and 28.24% on RoboCasa365, improving over the π₀.₅ baseline by 6.74 and 7.24 percentage points, respectively. Across three real-world manipulation tasks, it achieves 57.33% average success, a gain of 12.33 percentage points over the baseline. These results demonstrate improved task performance in both simulated and physical environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.