Investor-Specific Multi-Objective Preference Alignment for Personalized Portfolio Optimization
Abstract
Reinforcement learning for portfolio management typically outputs a single strategy for all investors, failing to reflect heterogeneous investor goals. To address this gap, we propose Investor-Type-Wise Multi-Objective Preference Alignment (IMOPA), a framework that accounts for diverse investor profiles and outputs strategies for different investors. Specifically, we integrate preference learning, bandit-based preference weight estimation, and policy optimization into IMOPA. IMOPA first evaluates candidate policies through LLM-generated multi-objective feedback on financial metrics together with an overall evaluation. Utilizing such information, a bandit-based mechanism estimates multi-objective preference weights and selects a candidate policy for each investor type. Finally, IMOPA employs Proximal Policy Optimization (PPO) conditioned on the learned multi-objective preferences to train personalized policies. Experiments on two datasets—S&P 500 and MidCap 100—demonstrate that our estimated preference weights reveal distinct financial trade-offs, leading to customized asset allocations across investor types. Furthermore, ablation studies show that omitting preference alignment narrows or collapses these behavioral differences. Overall, IMOPA enables interpretable, preference-aligned portfolio policies tailored to diverse investor profiles.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.