FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
Abstract
Trading-agent evaluation can fail at three successive levels: a ranking on realized history may not transfer reliably to alternative future paths, prospective returns may not exceed simple non-learning controls, and a favorable terminal outcome may persist after adaptation itself has stopped. We introduce FORESIGHT-9, a prospective and process-aware benchmark with nine auditable counterfactual worldlines branching from a common July 2026 information boundary. A com- mon contract fixes the market panel, cadence, constraints, costs, and account writes while preserving framework-specific research and proposal mechanisms. Across 36 long-horizon runs from two frameworks and two model backbones, Native-Policy Retrospective Replay yields a historical ranking whose median Spearman correlation with prospective rankings is −0.8; its winner remains first in only 1/9 worldlines. Prospectively, periodic equal weighting and a causal inverse- volatility rule each outperform 31 of 36 agent–worldline runs, showing that pos- itive agent returns need not establish incremental value from online adaptation. Finally, in a high-return WL8 run the live research library collapses while the declared ensemble persists and executed holdings converge toward equal weight: the run keeps earning after meaningful self-evolution has ceased. FORESIGHT- 9 therefore evaluates future robustness, incremental value over simple controls, and whether the intended adaptive mechanism remains functional. We release the worldlines, trajectories, audit traces, and deterministic regeneration scripts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.