RedCommerce: A Security Benchmark for Agentic Commerce Across the Full Lifecycle
Abstract
Agentic commerce requires agents to turn user requests into consequential commercial decisions. Determining a suitable action does not establish the evidence and authority needed to commit it. We introduce RedCommerce, a security benchmark organized around this Mandate–Commitment Gap. Its field-level formulation connects commercial choices to explicit attacker access, protected records, and executable outcome predicates. The benchmark contains 96 task packages and 829 variant-worlds across seven lifecycle stages and five role families. Four attacker configurations achieve 23.50–36.46% success against Haiku 4.5 under the original task predicates. On a matched 96-Case panel, saved adversarial submissions increase scored adverse outcomes by 21.53 percentage points relative to ordinary inputs (95% Case-bootstrap interval: 13.89–29.51). A post hoc sensitivity analysis excluding seven specification-sensitive Cases yields an 18.35-point increase. Repeated checks on eight selected attacks across four Cases show context-dependent benefits from object-binding and evidence-support assistance, while all 160 counterpart runs complete without scored deviation or escalation. Separate structured comparisons show that commitment decisions respond to the tested issuer and refund-scope conditions. These comparisons make the requirements supporting commercial actions experimentally inspectable beyond an aggregate attack-success rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.