Nash Equilibrium-Inspired Opponent Populations for Bayesian FPS Skill Diagnostics
Abstract
Estimating new player skill is a critical cold-start problem every multiplayer online game faces to deliver fair, competitive, and engaging gameplay. Current approaches over-compress skill into a single transitive rating based on a handful of matches. However, player abilities are multi-dimensional and often non-transitive: a player who dominates one aspect of gameplay may collapse on others. Onboarding bots, a popular solution, are static, easily recognizable by players, and largely uninformative. We introduce GAMBIT, a framework for Bayesian skill diagnostics that uses a calibration population of autonomous opponents and an adaptive testing loop. Rather than relying on a single transitive ladder, GAMBIT draws on Nash-equilibrium and population-based game-theoretic methods to bootstrap the population with hand-designed archetypes and progressively adds learned specialists to exploit weaknesses. This strategic diversity enables different opponents to probe different aspects of player behavior, becoming simultaneously both the challenge and the diagnostic instrument. GAMBIT operates in two stages. First, scripted and reinforcement-learned opponents are calibrated in a bot-versus-bot tournament to obtain relative difficulty estimates and uncertainty. Second, a new player competes a short sequence of bot opponents, where after each encounter a skill posterior is updated and the next opponent is adaptively selected to maximize uncertainty. This ensures the framework can be used against a continually evolving corpus of autonomous bots. In simulation, our results demonstrated 56.2% reduction in posterior uncertainty after just 6 opponents, and a final mean rating within 1.4% error margin from the calibrated reference. Furthermore, we analyzed the transferability and generalizability of these bots. Across 64 zero-shot transfers to unseen maps, the full policy stack completed without failure. Together, these results support strategically diverse autonomous opponents for uncertainty-aware cold-start skill estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.