acceptodds
Under review as a conference paper at ICLR 2027

Memory-Efficient Replay-Free Policy Evaluation for Mean-Field Games

Abstract

When the population size is large, a finite -agent game can be approximated by its mean-field limit, where the interaction among agents is represented through the population distribution. In finite-horizon evolutive mean-field games (MFGs), a common approach to computing a Nash equilibrium is to first generate the entire sequence of population distributions under a fixed policy and then fix this sequence as the external environment for policy evaluation. However, in many real-world systems, a simulator that can freely reproduce a prescribed distribution flow for subsequent policy evaluation is unavailable, which motivates a replay-free policy evaluation method. A straightforward alternative is to store the entire population history while a large finite population, approximated by the MFG, evolves forward to generate the distribution flow, and then perform policy evaluation after the forward evolution is completed. But this requires memory, which can be substantial when both the population size and the horizon are large. In this work, we first propose a Private Population Dutch Trace Algorithm, which recursively compresses the relevant history of each agent into a trace, eliminating the memory's dependence on . We further compress the collection of agent-wise traces and develop DoubleSWEET, which maintains state-wise traces instead, making the storage independent of and . Finally, we theoretically prove that the compressed state-wise traces in DoubleSWEET lead to similar policy evaluation as the private traces, for any fixed horizon, with an error upper bound of ; at the deterministic mean-field level, we further show that the same exact-expectation Dutch evaluator produces the same critic trajectory and learned value-function readout whether it is run live during population-flow generation or after that same flow has been frozen. In our experimental settings, DoubleSWEET achieves nearly the same policy-evaluation accuracy as the private-trace method while substantially reducing memory consumption; its memory usage does not grow with the population size or the episode horizon . At , the private-trace method requires as much memory as DoubleSWEET and the model-based projected evaluator requires approximately as much learner memory as DoubleSWEET-MC-LS while achieving the same relative RMSE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.