acceptodds
Under review as a conference paper at ICLR 2027

Conjunctive Amplification of Reward-Scale Miscalibration in Cooperative Multi-Event Reinforcement Learning

Abstract

Dense reward shaping is widely used in cooperative MARL, but in some multi-event tasks it can cause catastrophic performance loss. Reward-scale sensitivity is well known in deep RL; what remains unclear is why an ordinary scale mismatch becomes catastrophic only for particular task structures. We show that the missing link is conjunctive amplification. Miscalibrated rewards moderately reduce the acquisition rates of individual necessary events. When success requires all of those events, the local losses compound: all-events success closely tracks the product of per-event acquisition rates (). In a controlled Partial Relay environment, increasing dependence on the multi-event chain progressively narrows the empirically safe range across the tested cooperative reward weights, from all six values at to one at (225 runs; interaction ). Running reward calibration substantially mitigates this structure-amplified performance loss, with gains increasing from to as rises, and it generalizes across architectures and benchmarks, including a fourfold increase in Overcooked deliveries (). We further show why static tuning is unreliable: the effective reward scale changes during acquisition, so a good fixed scale can be identified only in hindsight. Finally, we introduce Event Order Consistency (EOC) as a low-cost behavioral descriptor of relay-family structure that can support calibration assessment before full training. These results support calibrating cooperative shaping when multi-event task structure can amplify small learning losses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.