Predict What Joint Actions Can Control: Goal-Progress Prediction for Post-Training Residual Selection
Abstract
Predictive models can guide action selection, but their usefulness depends on whether the prediction target captures differences that candidate actions can actually control. In cooperative multi-agent reinforcement learning, small coordinated interventions may have little influence on long-horizon outcomes, making it difficult to identify beneficial joint actions. We introduce Goal-Progress Estimation (GPE), a goal-progress predictive framework for post-training residual selection. Rather than reconstructing future states or predicting return to episode completion, GPE estimates short-horizon, candidate-dependent goal progress and uses these predictions to select coordinated residual actions around a frozen multi-agent policy. A conservative decision rule incorporates ensemble uncertainty and calibration, retaining the native policy as an exact fallback when predicted improvements are insufficiently reliable. On VMAS Transport, GPE improves offline candidate ranking and achieves substantially higher return than both native IPPO and a capacity-matched Direct Completion predictor. Guarded and greedy GPE achieve return gains of 1.228 and 1.802 over native IPPO, respectively, compared with 0.653 for Direct Completion. Further analyses highlight the role of prediction-target choice and conservative action selection, revealing a trade-off between return improvement and preservation of native-policy successes. Overall, our findings demonstrate the utility of short-horizon, action-controllable prediction for residual selection in cooperative multi-agent control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.