acceptodds
Under review as a conference paper at ICLR 2027

OW-CMORL: Online Window-Conditioned Multi-Objective Reinforcement Learning for Time-Varying Optimization

Abstract

This paper investigates non-stationary multi-objective reinforcement learning (MORL) for industrial control, where drifting exogenous operating conditions alter the Pareto frontier. Targeting this non-stationary online deployment scenario, we identify three practical obstacles: catastrophic forgetting, implicit context inference, and near-real-time constraints. We formalize this problem as a time-varying multi-objective contextual Markov decision process and derive three complementary theoretical upper-bounds bounding context embedding mismatch, local fine-tuning approximation, and long-run forgetting-induced performance degradation. We propose Online Window-Conditioned MORL (OW-CMORL): a context-indexed expert-bank stores restorable policy snapshots to mitigate catastrophic forgetting; a Transformer-based context-window encoder generates query embeddings from factor history and short-horizon forecasts for implicit unlabeled-condition representation; indexed retrieval fetches matched snapshots for lightweight local refinement, whose outputs refresh the online Pareto archive and avoid costly full retraining to satisfy near-real-time constraints. Experiments across four dynamic control systems under a unified evaluation protocol demonstrate that OW-CMORL preserves high-quality Pareto frontier while adapting to non-stationary operating conditions under near-real-time constraints. Open-sourced codes at: https://github.com/hoshinojyunn/OW-CMORL

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.