Urban-Centric Latent Multi-Task Offline Reinforcement Learning with Mask-Empowered Reward Learning
Abstract
In urban environments, human decision-making strategies—such as commuting routes, delivery logistics, and public transit planning—are often shaped by individual preferences and localized constraints, leading to suboptimal outcomes and leaving substantial room for improvement. Offline reinforcement learning (RL) offers a promising paradigm for improving urban decision-making by optimizing policies from pre-collected datasets, without requiring real-time interaction that could raise safety concerns. However, despite its potential, the application of offline RL to urban settings, particularly to realistic multi-task urban environments, remains largely underexplored due to the unique challenges inherent to urban decision-making. In this paper, we propose MORSE—Latent Multi-Task Offline RL with Mask-Empowered Reward Learning, a novel framework for optimizing diverse human strategies in urban environments. MORSE operates in a latent state space through Urban State Compression, improving learning efficiency and scalability. Moreover, it introduces a Mask-Empowered Reward Learning framework that integrates a novel trajectory masked autoencoder (TrajMAE) with a reward ensemble network, jointly trained in a semi-supervised manner, leveraging sparse labeled data from the target task alongside abundant unlabeled data from related tasks. The learned reward function is then used for data relabeling and augmentation. To mitigate the risk of over-reliance on relabeled data, MORSE further employs an Adaptive Knowledge Fusion strategy, enabling robust and generalizable policy learning in complex urban environments. Extensive experiments on real-world datasets demonstrate that MORSE significantly outperforms state-of-the-art baselines in learning diverse urban human strategies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.