acceptodds
Under review as a conference paper at ICLR 2027

The Decision Value of Partner Information in Cooperative Reinforcement Learning

Abstract

Cooperative agents often need information that is available only to their partners. Auxiliary objectives that predict a teammate’s behavior can improve internal representations, but prediction accuracy alone does not establish whether those representations improve coordination. We study this distinction by applying classical value of information to cooperative decisions and independently varying the availability and reward relevance of a partner’s private information. Experiments in an exactly solvable commitment task and modified Cooperative Reaching environments separate the value of observing a partner from the benefit of training an agent to predict it. With memory trained through time, recurrent policies without prediction objectives recover the available history value in the toy task and match a scripted-partner reference in the continuously observed reaching task. Auxiliary prediction can substantially increase the decodability of a reward-irrelevant partner variable while leaving return equivalent within a prespecified margin. When partner observations are restricted to an initial window, however, some co-learning runs remain below a prespecified task-return threshold at training budget. In this setting, a reward for behavior that makes the partner’s goal inferable increases the fraction of training runs exceeding the threshold from 42/60 to 60/60 and improves mean task return by 0.131 (95% confidence interval [0.085, 0.181]), excluding the auxiliary reward; a matched reward for progress toward the goal does the same, and given four times the budget the unrewarded runs cross the threshold too, so the gain is one of speed. Teammate prediction objectives yield no statistically resolved improvement; an extra input branch read by the policy helps, whether or not it is trained to predict the partner. These results distinguish representational gains from behavioral gains and show that, under restricted observation, encouraging informative partner behavior can speed up coordination where prediction objectives on their own do not improve coordination.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.