acceptodds
Under review as a conference paper at ICLR 2027

Exploitability-Guided Distillation of Interpretable Strategies

Abstract

Counterfactual regret minimization computes strong strategies for imperfect-information games, but as large tables of mixed strategies that people can neither audit nor execute. Deploying them in human-facing settings means distilling them into small decision trees, and the standard objective for that distillation is a distributional distance to the teacher. We show this objective is misaligned with strategic quality: at the same size, a leaf fit by imitation can be far more exploitable than one fit by exploitability. We propose exploitability-guided distillation. The tree is grown best-first under an exploitability split score, and leaf strategies are optimized directly against per-player best responses, using the separability of NashConv into independent per-player objectives; a single scalarized objective then traces the safety-fidelity frontier. On Leduc poker, best-response leaf fitting is 2.3x to 4.3x less exploitable than imitation fitting across budgets of 4 to 32 leaves, and at aligning the split score with the leaf objective cuts exploitability by 36 to 43 percent; the gain carries over to a five-rank variant. Our contributions are a bound that identifies the weights imitation misassigns to leaves, a distillation pipeline whose structure search and leaf solve share the exploitability objective, and an exact full-game evaluation showing that the leaf objective, with the teacher reduced to a uniform profile, returns the same Leduc structures at .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.