Hypervolume-Guided Iterative Reinforcement Learning for Multi-Objective Neural Combinatorial Optimization
Abstract
Multi-objective combinatorial optimization requires balancing competing objectives in routing, packing, and other discrete decision problems. Preference-conditioned neural methods commonly use scalarized rewards to train a single policy to construct solutions across different objective trade-offs. However, generating high-quality solution sets remains challenging because scalarized rewards overlook candidates' unique coverage within a group, while using each collected batch for only one policy update limits further quality improvements under a fixed sampling budget. To address this challenge, we propose Hypervolume-Guided Iterative Reinforcement Learning (HGI-RL), a training framework that combines marginal-HV advantage redistribution with repeated trajectory learning to improve solution-set quality. Specifically, exact marginal hypervolume contributions within each instance–preference candidate group redistribute positive advantages, incorporating candidates' unique coverage into the training signal. Moreover, the same trajectories and fixed redistributed advantages support multiple clipped proximal policy optimization (PPO) updates before new candidates are generated, allowing repeated learning from the collected solutions. To evaluate HGI-RL, experiments are conducted on multi-objective traveling salesman, capacitated vehicle routing, and knapsack problems. The results show that HGI-RL achieves competitive solution-set hypervolume.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.