Beyond High Prices: Auditing Information-Mediated Enforcement in Recurrent Pricing Agents
Abstract
High prices and correlated responses do not prove that pricing agents have learned collusive enforcement. We study this distinction in a continuous-price duopoly with noisy public signals and recurrent agents trained by proximal policy optimization. We use two benchmarks. A two-state trigger model shows when deviations are deterred by future losses. A Gaussian-monitoring model shows how noise weakens this threat. We audit the learned GRU policies using counterfactual price interventions and signal ablation. Across ten seeds and one million interactions, normalized price elevation drops from 0.43 in early training to 0.015 under informative monitoring, and to 0.019 after signal ablation. The paired difference is 0.0037, with a 95 percent bootstrap interval of [−0.0241, 0.0288]. Crucially, losses from forced price cuts hit immediately. They do not stem from adverse future values, and several individual deviations remain profitable. Rival responses are small under informative monitoring and vanish entirely under ablation. These results separate the illusion of coordination from actual dynamic enforcement. We find no evidence of real trigger enforcement in this setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.