Direct Pareto Optimization: Learning to rank for Multi-Objective Optimization
Abstract
Many real-world sequential design problems, such as molecule optimization and neural architecture search, require one to balance multiple conflicting objectives simultaneously. Prior approaches typically address this by decomposing the problem into a series of single-objective problems or optimizing auxiliary indicators such as hypervolume. While effective, these methods either do not directly optimize the quality of the solution set or are limited by the effectiveness of the indicators. In this work, we propose _Direct Pareto Optimization (DPaO)_, a framework that directly optimizes with respect to Pareto Dominance by casting the multi-objective sequential optimization task as a learning-to-rank problem. Specifically, we learn a scalar valued latent reward function from the _partial ordering_ induced by Pareto dominance, thereby transforming the _whole_ multi-objective problem into a single-objective one. We evaluate the proposed method in both synthetic and real-world tasks, where DPaO consistently ranks among top methods, regardless of the shape of the solution set or number of objectives. Furthermore, we empirically validate that the learned reward function aligns with the Pareto dominance relation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.