ActRight: Counterfactual Training for Causally-Grounded Vision-Language-Action Models
Abstract
Vision–Language–Action (VLA) models achieve strong performance in robotic manipulation, but existing approaches optimize task success without ensuring that actions are grounded in language semantics. As a result, models often rely on spurious visual correlations, leading to failures under distribution shifts and novel vision–language compositions. We propose ActRight, a training framework that enforces counterfactual causal constraints for cross-modal decision making, with three key technical contributions. First, we impose a structured V→L→A informa- tion flow by introducing an intermediate language representation as the sole input to the policy. Second, we introduce counterfactual objectives that constrain model behavior under interventions on vision and language, enforcing uncertainty under mismatched inputs and sensitivity to changes in language semantics. Third, we de- velop an adversarial mechanism that actively constructs challenging counterfactual scenarios to expose and penalize violations of causal consistency during training. Experiments on the LIBERO and CALVIN benchmarks show that ActRight consis- tently outperforms existing VLA methods in robustness and generalization, while improving sensitivity to language conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.