acceptodds
Under review as a conference paper at ICLR 2027

Offline Equilibrium Finding in Extensive-form Games: Datasets, Methods, and Analysis

Abstract

Offline reinforcement learning (Offline RL) has emerged as a promising paradigm for solving real-world decision-making problems using pre-collected datasets. However, its counterpart for extensive-form games (EFGs) has received little attention. To fill this gap, we introduce the *offline equilibrium finding (Offline EF)* paradigm, which aims to compute equilibrium strategies of an EFG from a fixed offline dataset without interacting with the game. Offline EF is harder than offline RL, and the difference lies in the solution concept: offline RL seeks an optimal policy, whereas offline EF seeks an equilibrium. In single-agent offline RL, deciding whether one action is better than another only requires comparing the two in the data. An equilibrium, on the other hand, is a claim about every player's best response: no one can gain by switching to any other strategy. A dataset that records only how the players actually played says nothing about these deviations, and so cannot certify an equilibrium. Existing equilibrium-finding algorithms sidestep this by querying the game for best responses, which a fixed dataset cannot answer. Additionally, there are no public datasets or evaluation protocols for this problem, which leaves methods incomparable and progress hard to measure. To address these issues, we first construct a suite of offline datasets over a range of EFGs, spanning random, expert, learning, and hybrid data-generating strategies. We then propose BOMB, a framework that combines behavior cloning with a model-based method and allows any online equilibrium-finding algorithm (e.g., CFR, PSRO) to be used in the offline setting. We identify the sufficient conditions under which each component of BOMB is guaranteed to converge to an equilibrium, and show that BOMB, with its optimal mixing weight, is at least as good as the better of its two components. Extensive experiments show that BOMB consistently outperforms offline RL baselines and computes approximate Nash and coarse correlated equilibria in both two-player and multi-player games.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.