acceptodds
Under review as a conference paper at ICLR 2027

Directed Laplacian Representation Learning for Non-Reversible Markov Chains in RL

Abstract

This paper presents a method for Laplacian representation learning in directed Markov chains induced by a fixed reinforcement-learning policy, using the Chung Laplacian as the spectral target. Prior approaches estimate this target from discounted forward transitions, which is exact on reversible chains or at a single step. But a fixed policy generically induces a non-reversible chain, and learning the representation from its stationary transitions poses two challenges, as the energy and orthonormality constraints depend on the unknown stationary distribution, while the forward multi-step sampling can distort the desired representation. Learning in weighted coordinates removes the explicit dependence on the stationary distribution, while a discounting scheme based on the reversible component of the dynamics preserves the target eigenvectors. The method combines forward and time-reversed transitions to construct an unbiased estimator of this energy under stationary sampling. Coupling this estimator with an augmented Lagrangian objective yields local stability guarantees for spectral recovery under a simple-spectrum assumption. Experiments show that the method matches competing approaches in reversible environments and substantially outperforms them in recovering the eigenvectors of the Chung Laplacian in non-reversible settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.