acceptodds
Under review as a conference paper at ICLR 2027

Behavioral Cloning for the Action Policy of a Customer Simulator: Structured Scores, Metric-Aware Decisions

Abstract

Simulating a customer journey requires predicting not only which product a shopper engages with, but whether they search, refine, navigate, convert, or stop interacting. Recent shopping simulators build this on language models, whose post-training targets helpful interaction rather than a cohort's action propensities, and task adaptation does not by itself guarantee recovering them—prior work reports homogenised, over-agreeable behavior when they stand in for people (Zhou et al., 2026; Sharma et al., 2024), a distortion that would land on the conversion and termination rates a simulator exists to produce. We factor a simulator step by the chain rule into behavior selection and semantic realisation, and ground the first in logged behavior by cloning. We isolate a journey-spanning action policy covering query formulation, item engagement, conversion, navigation and termination; 40.4% of test steps carry an action no standard item-ID next-item objective represents. We introduce StructBC, a compact causally masked Transformer score model encoding product-slate structure, short-range temporal dependence, page-local state and multi-scale behavioral context—a composition that lifts a recurrent and an attentive backbone alike. Our corpus is 452,296 real shopping episodes from 6,824 shoppers, split customer-disjointly into train, validation and test, over a state audited channel by channel for time direction (the past-only state: every shopper-specific channel is prefix-valid, bar two disclosed non-shopper transforms). On it the 3.4M-parameter policy reaches 0.5065 macro-F1 and 59.0% top-1 accuracy over ten actions, against 0.2784 for the strongest count-based baseline on the same split. Over the nine directly observed actions the score model alone reaches 0.5375 under the same argmax decoding, against 0.5278 for an adapted NARM encoder, and it also outperforms, under a fixed single-GPU compute envelope, a LoRA fine-tune of a 1.5B instruction model on our corpus and label. Under the same recipe extended to all ten actions, the LLM predicts termination on 0.15% of matched test steps against 4.61% observed—a prevalence error our own scorer reproduces under a comparably data-starved, one-epoch control. A nine-parameter metric-aware decision layer, fit post-hoc with no retraining of the score model, takes the complete policy to 0.5488 over the nine observed actions (+0.0113 of that from the layer), and contributes +0.0152 over all ten labels. It also cuts the total-variation distance between predicted and observed action distributions from 0.1055 to 0.0498. We evaluate the action-selection factor open-loop: the discrete policy of a customer simulator, composable with the semantic realisation we do not evaluate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.