The Interface Is the Policy: Output Typing, Calibration, and Confidence-Gated Routing in LLM Trading Agents
Abstract
Large language model (LLM) trading agents are evaluated as if the model were the policy: frameworks differ in memory, reflection, and multi-agent structure, but nearly all terminate in a generated decision that a parser reduces to buy, sell, or hold. We argue that this final step, the output interface, is a design axis whose effect on decision quality has never been isolated. We formalise an interface ladder of four rungs that hold the model, prompt, and market context fixed while varying only what the model is asked to emit: regex-parsed free text (L0), a constrained enum (L1), an enum with verbalized confidence (L2), and a typed probability vector over next-day direction (L3). The ladder exposes a structural fact that prior benchmarks obscure: rungs L0–L2 can only be mapped to a full-size long, short, or flat position, whereas L3 admits fractional sizing, so any advantage of “primitivized” outputs must flow through the probability-to-position map rather than the label. We instantiate the ladder on InvestorBench, an open benchmark of seven equities and cryptocurrencies, with open-weight backbones from 7B to 72B; audit whether the confidence emitted by L2 and L3 is actionable, meaning calibrated, discriminative, and stable across assets; and sweep the two-threshold “auto-execute above τ_hi, escalate below τ_lo” routing rule against always-execute. Asking the same model for a different output type changes its direction on roughly half of all days (agreement with the enum rung 0.52, κ = 0.04), so the interface is part of the policy, yet no rung has detectably better risk-adjusted return under sign-only sizing. Mapping L3’s probability vector to a fractional position cuts volatility to 0.25× and drawdown to 0.21× at equal or better Sharpe, which is where typed outputs earn their keep. The emitted confidence is honest after temperature scaling (ECE 0.117 → 0.047) but not discriminative (AUROC = 0.530), so no confidence-gated routing policy dominates always-execute. Code, prompts, decision logs, and the harness are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.