Taylor Representations for Model-Free RL in Networked MDPs
Abstract
In Networked Markov Decision Processes, transition dynamics are often unknown and the state-action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate -functions, but a naive order- expansion over agents requires coefficients. We justify these expansions under smooth expected future local rewards with controlled derivatives. Under this condition, finite-speed information propagation and discounting imply that local-critic Taylor coefficients decay exponentially with the graph distance to the farthest agent involved. Discarding distant-agent coefficients and marginalizing then yield scalable local Taylor representations with a bound controlled by graph locality. Building on these representations, we propose a scalable model-free actor-critic algorithm, establishing finite-sample critic and near-stationarity guarantees for a linear LSTD critic. We then introduce a more expressive neural TD parameterization. Unlike prior constructive spectral methods, our approach covers settings without access to a known local dynamics map, such as hidden switched linear-quadratic regulation. Across three control benchmarks, our method matches or outperforms spectral baselines while scaling efficiently to large graphs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.