acceptodds
Under review as a conference paper at ICLR 2027

ADX-Bench: A Benchmark for LLMs in Advertising Forecasting, Diagnosis, and Decision Making

Abstract

Large language models (LLMs) are increasingly used for domain-specific reasoning and decision support, yet their capabilities in real-world advertising remain insufficiently evaluated. Advertising decision making is inherently sequential: models must anticipate future outcomes, diagnose merchant-specific performance issues, select actions under explicit business constraints, and determine whether committed decisions should be revised as new information becomes available. Existing evaluations typically abstract these capabilities into isolated tasks and focus primarily on task-level correctness, providing limited insight into whether model behavior is grounded, operationally valid, and reliable throughout the decision process. To address these challenges, we introduce ADX-Bench, a decision-centered benchmark for evaluating LLMs in real-world advertising. ADX-Bench comprises three complementary benchmarks constructed around a shared data lineage. Action-Conditioned Outcome Forecasting (AOF) evaluates multivariate multi-step forecasting from historical advertising trajectories, campaign context, and observed actions. Evidence-Grounded Diagnosis (EGD) evaluates taxonomy-constrained root-cause identification, merchant-specific evidence grounding, and intervention recommendation. Constraint-Aware Action Selection (CAS) evaluates action selection under business objectives and operational constraints, together with post-commitment reasoning when additional evidence and critiques are revealed. Built from real-world operational advertising data, ADX-Bench preserves the sequential dependencies and evolving information states of advertising decisions while enabling independent evaluation of predictive, diagnostic, and decision-making capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.