acceptodds
Under review as a conference paper at ICLR 2027

SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long Horizon Manipulation

Abstract

At every re-observation point in a hierarchical Vision-Language-Action (VLA) system, two interface decisions must be made. The first is when to terminate the current subtask. The second is how far to execute the proposed action chunk. These decisions are mutually dependent. The useful stopping point depends on the proposed executor behavior, while the useful execution length depends on subtask progress. Many architectures nevertheless evaluate them with separate decision pathways, which leaves an information asymmetry that neither module can overcome alone. We present SparkVLA, a stop-aware hierarchical VLA that implements this coupled decision by formulating both choices as a single ranking. In this ranking, Stop competes against every action-prefix length in a unified candidate set, and the system selects the highest-scoring option, avoiding a separately tuned probability threshold and requiring only offline ordinal preferences. An Anchor-Conditioned Context Encoding module caches a history-aware subtask anchor that encodes onset-state memory and goal semantics, and guides visual-token pruning toward task-relevant regions. A Stop-Aware Action-Prefix Selection} head scores all candidates via full self-attention at chunk boundaries for efficiency. On RoboCerebra, SparkVLA achieves 36.1% complete-task success and 47.12% transition completion, surpassing the official hierarchical baseline by 30.10 percentage points on complete-task success. Real-robot experiments on multi-step tasks further validate these gains on physical hardware.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.