LOCM: Behavioral Refinement of Reward-Induced Trajectory Equivalence in Multi-Agent Reinforcement Learning
Abstract
In cooperative multi-agent reinforcement learning, different policies may achieve similar cumulative returns while exhibiting distinct execution behaviors. We interpret this return degeneracy as a structural consequence of compressing high-dimensional joint trajectories into a scalar task return, thereby inducing behaviorally heterogeneous return-equivalent sets. Our theoretical analysis indicates that scalar task returns can locally induce high-dimensional sets of return-equivalent trajectories, while appropriately chosen trajectory-level behavioral representations can further distinguish behavioral variations that remain unresolved by the original task return. Based on this view, we propose LOCM (LLM-Calibrated Observable Coordination Modeling), a trajectory-aware behavioral refinement framework that further refines execution behavior within an -optimal task-performance set rather than replacing the original task objective. LOCM combines interpretable joint-action behavioral coordinates, phase-adaptive weighting through an LLM-assisted semantic interface and numerical adapter, and an outcome anchor that constrains task-performance degradation. Experiments on power distribution network control and traffic-control tasks show persistent behavioral differences among near-return trajectories, identify locally reducible execution overhead, and demonstrate that LOCM reduces action variation across multiple MARL backbones while maintaining competitive task performance. These results support task-preserving behavioral refinement as a principled mechanism for mitigating reward-induced execution ambiguity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.