acceptodds
Under review as a conference paper at ICLR 2027

When to Rely on a Partner? Reliability based Arbitration for Zero Shot Cooperation with Decision Process Impairment

Abstract

Training agents that cooperate with humans requires adapting to diverse partner behaviors, which cannot be fully anticipated before interaction. Zero shot coordination (ZSC) addresses this by exposing agents to a wide range of partner behaviors during training, covering differences in skill level, strategy, preference, and action stochasticity. Such coverage, however, presumes partners whose decision processes are intact. Agents deployed as assistive or collaborative partners must also work with people whose decision processes are impaired, as in cognitive and psychiatric conditions. Such impairment does not lower competence uniformly: it selectively affects behavior during the task, so the same partner may act competently in one situation and fail in another. A fixed cooperation policy is therefore not enough; the agent has to keep reassessing how much to rely on the partner as the interaction unfolds. Adapting behavior across different situations is a fundamental aspect of human cognitive control. Reliability based arbitration control has been identified as a mechanism that supports such adaptation by adjusting the influence of competing strategies according to their estimated reliability. Inspired by this principle, we propose Reliability based Dual Specialist Arbitration (RDSA). RDSA decomposes cooperation into two strategies, each based on a different assumption about the partner's reliability: a Partner-Reliant Specialist that relies on the partner to handle part of the task, and a Self-Reliant Specialist that completes the task on its own. During interaction, RDSA estimates each specialist's reliability and dynamically adjusts their relative influence on action selection. To evaluate RDSA, we constructed simulated partners with decision process impairments by applying targeted perturbations to perception, valuation, and action selection. Experiments showed that these partners exhibited behavioral variation distinct from skill level, with the differences depending on the situation. RDSA outperformed the best performing ZSC baseline and the best static mixture in most evaluated conditions, including conditions in which neither specialist outperformed the best performing ZSC baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.