RAFT: Reasoning Aligned Fine Tuning Improves Robustness of VLA Post Training
Abstract
Adapting a generalist vision-language-action (VLA) policy to a target task with limited demonstrations can improve nominal task success while reducing robustness to variations in appearance, objects, and instructions. We introduce reasoning- aligned fine-tuning (RAFT), a two-stage post-training procedure: RAFT first trains low-rank adapters on task-relevant reasoning annotations, then freezes them while fine-tuning the policy on action demonstrations. On a physical DROID robot, RAFT achieves about 75% out-of-distribution success, compared with about 28% for plain fine-tuning; on a multi-stage bimanual YAM task, it reaches 70% mean success, compared with 36% for the strongest baseline. RAFT also improves success under visual, object, and instruction shifts in simulated pick-and-place and drawer-opening tasks. Compared with plain fine-tuning, RAFT policies place more attention on annotated task regions and yield higher linear-probe accuracy for object identity in later layers, and a broader reasoning-data mixture further improves downstream success across demonstration budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.