AlphaDreamer: Model-Predictive Prefix Planning for Formulaic Alpha Discovery
Abstract
Discovering predictive alpha factors is fundamental to quantitative trading. However, evaluating a formula’s predictive quality requires completing its construction, leaving intermediate decisions without direct performance feedback. We introduce AlphaDreamer, a model-based reinforcement learning framework that learns predictive models from evaluated formulas to guide construction through prefix-level planning. Its Model-Predictive Prefix Planner (MPPP) couples exact grammar transitions with learned predictions of prefix representations and completion quality. A learned reference model directly scores candidate prefixes to complement predictive lookahead in action selection. During prefix planning, the planner executes only the first action of each selected plan before replanning from the updated prefix. A shared policy learns both to generate formulas from scratch and to complete planner-provided prefixes. Empirical results show that AlphaDreamer outperforms the evaluated baselines in predictive accuracy and portfolio performance. Prefix-level planning improves both discovery quality and evaluation efficiency, while the full planner consistently outperforms direct prefix scoring alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.