ACTION–EFFECT CONTRACTS: BRIDGING TASK INTENT AND PHYSICAL EFFECTS IN AGENTIC ROBOTS
Abstract
Vision-language-action (VLA) models map language instructions and visual ob- servations directly to continuous robot actions. Real-world robot tasks are often long-horizon and compositional, requiring agentic robot systems to plan and ex- ecute sequences of interdependent subtasks. In mainstream designs, a high-level agent intervenes through task-flow decisions such as subtask invocation, stopping, replanning, and human assistance, while the low-level VLA determines the specific effects of task execution. This architecture leaves the agent without direct involve- ment in action execution. To increase the agent’s participation during execution, we introduce Action–Effect Contracts, a semantic Agent–VLA interface that lets the high-level agent specify the physical effect required at the current task stage. At each decision point, a frozen VLA proposes multiple short action candidates, and an effect ranker selects the candidate most likely to realize the requested effect. The ranker jointly conditions on the current image, robot state, effect request, and candidate action through candidate-dependent spatial attention. Same-state counterfactual rollouts provide candidate-level matching labels and within-group ranking supervision. We instantiate the high-level agent as a fine-tuned Qwen3-VL- 8B-Instruct model that generates online effect requests, and use a frozen π0.5 action policy for short-horizon manipulation in LIBERO. Across nine LIBERO-90 tasks, our method improves requested-effect matching in all four stages and raises task success from 42.06% to 51.27%. In four real-world manipulation tasks performed by a robot arm, success increases from 48.33% to 58.33%. These results show that high-level effect requests can effectively guide VLA action execution and improve task completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.