RetailAgent: Auditing Systematic Timing Errors in Self-Conditioned LLM Agents
Abstract
Sequential LLM agents can fail in ways that are directional rather than noise-like. We introduce RetailAgent, a controlled multimodal, state-conditioned framework that repeatedly chooses long or flat from an anonymized intraday price history, its chart representation, and permitted account state. RetailAgent provides a sparse price-and-state interface for auditing sequential decisions; it neither represents a retail population nor simulates a full market. We evaluate each decision trajectory using exposure-matched within-stock timing, which tests alignment with subsequently revealed return labels rather than executable portfolio performance. Across text, chart, joint text–chart, and account-state conditions, all estimates are negative; the principal 10-minute text, chart, and multimodal arms are , , and bps per stock-day. This wrong-signed timing is directional: the complementary schedule reverses its score, while linear projection on conventional price models leaves substantial residual timing. Narrative self-conditioning further reorganizes the trajectory: stronger wrong-signed timing coincides with fewer position changes in the tested memo conditions. We interpret these findings as structured sequential failure and as motivation for future counter-policy and leader–follower analyses, not as evidence of profitable trading, market impact, or human retail equivalence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.