acceptodds
Under review as a conference paper at ICLR 2027

Plan Before You Trade: Inference-Time Optimization for RL Trading Agents

Abstract

Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for policy adjustments using price forecasts at inference time. We propose FinPILOT (Financial Plugin Inference-time Learning for Optimal Trading), a plugin inference-time policy optimization framework inspired by Model Predictive Control (MPC). Under the price-taker assumption, portfolio actions do not affect future market prices, allowing a multi-step price forecaster to provide imagined market trajectories without iterative action-conditioned rollouts. At each decision step, FinPILOT uses the forecaster's predicted price trajectory to construct an allocation-based imagined return objective, and optimizes the pre-trained policy at inference-time before executing one step of the trade. Our framework is compatible with any pre-trained policy and adapts the policy using the forecaster's predictions without modifying the original RL training pipeline. Evaluated with five pre-trained policies (PPO, SAC, A2C, TD3, DDPG) across DJ30 equities and foreign exchange (FX) over three rolling held-out test years (30 settings in total), FinPILOT improves Total Return in 28 of 30 settings and Sharpe ratio in 27 of 30, with gains spanning both stochastic and deterministic policies. Direct single- and multi-period portfolio optimizers supplied with the same forecasts do not reproduce these aggregate gains under our evaluation protocol. Further, using controlled forecasts at calibrated quality levels, we show that both the magnitude and consistency of FinPILOT's performance increase as forecast error is reduced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.