Vision-Language-Action Models are Inherent Verifiers: Test-Time-Scaling via Likelihood-Based Visual Prompting
Abstract
Test-time scaling (TTS) has recently emerged as a powerful paradigm for improving Vision-Language-Action (VLA) models without requiring expensive fine-tuning. However, existing TTS approaches typically rely on training external verifiers or are strictly limited to autoregressive action decoders. To address these limitations, we introduce Likelihood-Based Visual Prompting (LVP), a training-free, verifier-free test-time scaling framework and reveal that generalist VLA policies can serve as their own inherent verifiers. LVP first projects action candidates directly into the 2D image space via perspective-aware action projection, overlaying future trajectory and gripper poses as structured visual prompts. We then reframe action verification as an evaluative Visual Question Answering (VQA) task. By prompting the VLA to evaluate the quality of the visually prompted action via a multiple-choice question (e.g., grading the action from "Poor" to "Excellent"), we leverage the model's internal conditional token likelihoods to compute an expected action quality score for candidate ranking. Furthermore, we introduce an attention-guided token pruning mechanism to alleviate the computational overhead of processing multi-sample visual prompts. By leveraging the visual attention maps generated during the action proposal stage, LVP dynamically discards task-irrelevant visual tokens during the verification phase. In the experiments, LVP improves autoregressive and diffusion/flow-matching based VLA models across diverse simulation and real-world benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.