acceptodds
Under review as a conference paper at ICLR 2027

Efficient Multi-Agent Reinforcement Learning under Sparse Binary Team Rewards with Few Suboptimal Prior Trajectories

Abstract

Learning efficient cooperative policies under sparse binary team rewards is particularly challenging when only a few suboptimal prior trajectories are available. Binary terminal feedback collapses trajectories of different qualities into identical outcomes, obscuring value estimation, while representative methods struggle to recover these hidden distinctions. We therefore propose TRACE, an online MARL framework that restores within-outcome reward resolution under binary feedback and converts the resulting trajectory-level supervision into discriminative transition-level learning signals. TRACE requires only the meaning of the original sparse binary reward and available terminal-state variables. Given this limited information and the proposed evaluator template, a fixed terminal-state graph evaluator can be instantiated either manually or with LLM assistance. The evaluator provides discriminative trajectory-quality scores within each outcome class while preserving success-failure separation. Uniform discount-aware relabeling then propagates this quality-derived terminal reward to preceding transitions without changing the induced policy ordering. Progressive prior-guided co-training integrates limited suboptimal priors with online experience while gradually reducing prior dependence. Across discrete-action SMAC v1/v2 and continuous-action MPE benchmarks, TRACE achieves the best overall training efficiency and performance. Further results show improved value discrimination, and robustness to prior quantity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.