DASH: Diagnostic-Guided Quality-Diversity Search for Financial ML Agents
Abstract
Automated LLM research agents iteratively propose, evaluate, and refine machine-learning configurations using empirical feedback. In financial forecasting, low signal-to-noise ratios and nonstationarity make apparent improvements uncertain and potentially transient, allowing search to reinforce spurious or regime-specific gains. We introduce DASH (Diagnostic Agentic Search with Hypotheses), a diagnostic-guided search framework that evaluates configurations across temporal regimes to construct condition-specific, multi-axis behavioral profiles. Before refining a candidate, the agent formulates hypotheses with pre-specified expected outcomes and excludes those refuted by computed evidence. A quality-diversity archive organizes evaluated states by behavioral descriptors while separately accounting for predictive strength and redundancy when retaining and selecting alternatives. Empirically, DASH achieves strong held-out OOD predictive and portfolio performance under matched search budgets. In joint factor–model search mode, it improves IC over the strongest evaluated agentic baseline by 8.7% on NASDAQ-100 and 38.2% on S&P 500, while achieving the highest portfolio IR on both markets. Factor-only analyses show a high concentration of strong OOD factors with low predictive redundancy in the discovered factor pools. In a separate NASDAQ-100 ablation, replacing externally computed evidence with LLM-only judgments consistently degrades predictive performance across two backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.