ObjectBinder-VLA: Permutation-Consistent Multi-View Grounding and Intent-Conditioned Dynamics
Abstract
Vision-language-action (VLA) policies commonly condition control on global visual-language representations that do not explicitly encode task-relevant object support, visibility, multiplicity, or cross-view correspondence. We introduce ObjectBinder-VLA, a predictive multi-view object-state interface that preserves global VLA context while exposing a compact, measurable state for control. For each task-conditioned slot and view, a shared Object Binder predicts a categorical distribution over 64 image patches and a NO_OBJECT state, together with independent visibility. A fine-tuned SAM3.1 teacher provides training-only pseudo masks; Hungarian set matching avoids arbitrary identities for exchangeable instances, and teacher-aware duplicate regularization discourages slot collapse. An Intent Query and Object Dynamics Model predict this state at on LIBERO and on RoboTwin. Detached future distributions are projected into view-slot tokens for a flow-matching action expert. Stop-gradient boundaries isolate direct updates to the task-specific perception and prediction modules, while the shared Qwen backbone remains jointly optimized. Held-out diagnostics demonstrate spatial grounding and visibility prediction, with no failures across 100 target-order checks of the set objective. The final 50K model achieves 98.10% success (1,962/2,000) on Original LIBERO. These results indicate that explicit object binding and intent-conditioned dynamics can provide auditable intermediate states without replacing global policy context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.