acceptodds
Under review as a conference paper at ICLR 2027

Universe-as-Code: Executable Asset Selection for Portfolio Reinforcement Learning

Abstract

Portfolio reinforcement learning depends on both the assets available to a policy and the reward used to train it. We propose UCAPS (Universe-as-Code Agent Portfolio System), which learns executable upstream design artefacts for these coupled decisions. Universe-as-Code (UaC) generates multiple asset-selection programs, validates them in a sandbox, and combines their outputs through deterministic per-ticker voting. Dual-Artefact Co-Evolution (DACE) evaluates universe–reward pairs through downstream policy training and uses successful outputs to update separate Selector and Coder adapters. Its Composite Robust Score (CoRS) aggregates multi-seed feedback on historical performance, drawdown, temporal dispersion, and program complexity. On a 30-ticker multi-asset pool, UCAPS achieves 44.1% mean Test cumulative return and 1.31 Sharpe across four Selector backbones and 80 policy runs, net of transaction costs. The mean return exceeds the strongest reported same-pool large language model (LLM) agent adaptation by 10.5 percentage points. Price-matched classical baselines supply a common data reference, and eight matched selection controls isolate the selection rule. Under a fixed reference reward, the Qwen2.5-7B and Llama-3.1-8B Selectors exceed the strongest alternative by 20.6 and 11.1 percentage points. These results establish executable universe selection as an effective upstream optimization target for portfolio reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.