acceptodds
Under review as a conference paper at ICLR 2027

From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

Abstract

Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often cover periods before the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for the ability to select stocks, since positive performance may come from market beta, style exposure, or favorable regimes rather than stock-selection alpha. We introduce Knowing-To-Doing Financial Benchmark (KTD-Fin), an end-to-end trading benchmark based on price data that addresses both issues. To mitigate knowledge leakage when backtests cover periods before LLMs' knowledge cutoffs, KTD-Fin applies masking at the data layer to anonymize key identifiers and calendar information consistently across prompts and tools, limiting the cues agents can use to retrieve memorized market information while enabling controlled comparisons of identifier access. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. This finer-grained decomposition allows stock-selection ability to be assessed separately from market and style exposures that can confound evaluations based on raw returns. Across ten frontier LLM agents evaluated on the Chinese A-share CSI300 over a 2024–2026 trading window, masking changes agent rationales, pushing them away from narratives about specific companies and toward factor reasoning under anonymized identifiers. Attribution analysis further shows that LLM agents' cumulative returns under identifier masking are largely explained by market and style exposure, with negative stock-selection alpha for nine of ten agents under the full nine-factor specification. Cross-model replications reproduce behavioral differences; leave-one-factor-out attribution preserves closely aligned rankings. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also which sources drive those returns. We provide KTD-Fin as a reproducible template for evaluating LLM trading agents with leakage controls and return attribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.