OffRoad-VLA: Terrain‑Aware Action Tokens and Affordance Reasoning for Off-Road VLA Navigation
Abstract
Vision-Language-Action (VLA) models have recently gained significant attention in autonomous driving for their ability to leverage world knowledge for navigation. However, existing VLA approaches primarily target on-road environments, where physical actions assume structured road surfaces and traffic semantics. We present OffRoad-VLA, an end-to-end autoregressive VLA framework that unifies terrain-aware reasoning and physical action generation for off-road navigation. We introduce a terrain-aware physical action codebook that jointly discretizes vehicle motion with local terrain properties, including slope, roughness, and traversability, enabling the model to generate kinematically feasible and terrain-adaptive trajectories. To support efficient reasoning, we employ dual thinking modes that produce direct actions for simple scenes and adaptive thinking using chain-of-thought (CoT) reasoning for complex scenario, explicitly connecting terrain conditions and obstacles to driving actions. We further apply Group Relative Policy Optimization (GRPO) with off-road safety rewards to improve planning while penalizing traversability violations, unsafe terrain interactions, encouraging progress, and smooth motion. Experiments on RELLIS-3D and TartanDrive, together with zero-shot evaluation on the unseen GOOSE dataset, demonstrate improved long-horizon trajectory prediction, terrain classification, and safety over existing approaches. Closed-loop evaluation in the MAVS physics-based simulator further demonstrates robust real-time navigation across diverse off-road conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.