acceptodds
Under review as a conference paper at ICLR 2027

Which Environments Teach Cooperation? Guarantees, Feedback, and Competence

Abstract

Choosing environments for independently trained agents involves two different questions: are their near-optimal self-play solutions compatible, and can the learners reach those solutions? We study this distinction while requiring competence on a fixed task distribution. For sequential protocols with shared preferences and action-independent paths, an exact balanced partition of diagnostic losses characterizes incompatible near-optimal pairs. Conflicting one-step sources instead induce signed exposures; conflicting sequential sources need not admit the same representation. We then show that observable failure locations lead to a polynomial singleton-exposure objective for a correction learner, even when a terminal-reward estimator requires much more data. Experiments separate solution-set guarantees from learned performance. On invariant tasks with delayed rewards, training on the competence distribution is a strong default. In small action-steering tasks, controls with matched target information reveal that the advantage of margin selection depends on the competence floor and nominal tie-breaking rule. Additional mixtures selected independently of the margin show a positive association within positive-margin regimes. The resulting framework gives conditional certificates and identifies which environmental, feedback, and competence assumptions are needed to interpret them; it does not make a large margin a universal curriculum objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.