Universe-as-Code: Executable Asset Selection for Portfolio Reinforcement Learning
Abstract
Portfolio reinforcement learning depends on both the assets available to a policy and the reward used to train it. We propose UCAPS (Universe-as-Code Agent Portfolio System), which learns executable upstream design artefacts for these coupled decisions. Universe-as-Code (UaC) generates multiple asset-selection programs, validates them in a sandbox, and combines their outputs through deterministic per-ticker voting. Dual-Artefact Co-Evolution (DACE) evaluates universe–reward pairs through downstream policy training and uses successful outputs to update separate Selector and Coder adapters. Its Composite Robust Score (CoRS) aggregates multi-seed feedback on historical performance, drawdown, temporal dispersion, and program complexity. On a 30-ticker multi-asset pool, UCAPS achieves 44.1% mean Test cumulative return and 1.31 Sharpe across four Selector backbones and 80 policy runs, net of transaction costs. The mean return exceeds the strongest reported same-pool large language model (LLM) agent adaptation by 10.5 percentage points. Price-matched classical baselines supply a common data reference, and eight matched selection controls isolate the selection rule. Under a fixed reference reward, the Qwen2.5-7B and Llama-3.1-8B Selectors exceed the strongest alternative by 20.6 and 11.1 percentage points. These results establish executable universe selection as an effective upstream optimization target for portfolio reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.