MA-RL-VLA: Counterfactual Credit-Aware PPO for Multi-Agent Reinforcement Fine-Tuning of Vision-Language-Action Models
Abstract
Reinforcement learning (RL) has proven effective for fine-tuning single-agent vision-language-action (VLA) models, yet extending this paradigm to multi-agent vision-language navigation (MA-VLN) remains largely unexplored. We present MA-RL-VLA, a centralized-training / decentralized-execution (CTDE) framework whose core design is a mathematically consistent integration of counterfactual credit into PPO for VLA-scale actors. The counterfactual baseline is identical in form to COMA's policy-marginalized baseline, and because one actor pass already yields the full distribution over the four discrete navigation actions, it is computed by exact enumeration at the cost of four small-critic evaluations per agent step. The detached, clipped credit weight is folded inside both PPO branches as a combined advantage , preserving PPO's pessimistic clipping under negative credit and thereby suppressing free-riding. The reward separates a policy-invariant potential-based progress term (fixed, never tuned) from auxiliary coordination terms with fixed weights selected on training-split tuning scenes. On CoNavBench (Habitat, ) with 8 matched seeds, MA-RL-VLA improves Team Success Rate by points over behavior cloning (95% CI ) and points over standard COMA-style credit (; exact paired permutation ), and cuts exploration redundancy by . Controlled ablations attribute the gain to the credit signal (), the objective form ( over the legacy outer-weight surrogate), and the exact estimator. Our evidence is confined to one simulated benchmark and one VLA backbone; we state precisely which guarantees do and do not transfer to the deployed update.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.