From Generation to Decision: Proper Scoring Rules for Type-Safe User Action Prediction
Abstract
User simulation must reproduce behavior on interfaces that mix discrete choices with open-ended text. Existing language-model simulators generate action strings for both, coupling finite choices to token decoding. We introduce CADET (Calibrated Decisions over Types), a hybrid simulator that matches its computation to the interaction. A pre-trained LM encoder feeds a typed decision head that predicts joint operation-subtype probabilities in one forward pass. An element selector resolves the target, and the predicted operation gates a language generator for text-bearing actions. Proper scoring rules train these probabilities to recover behavioral distributions beyond the most likely action. We evaluate search and category-based browse on two e-commerce benchmarks. On one benchmark, CADET-Brier reduces expected calibration error by 27–58% across seven frozen backbones relative to the best regularized exact-match baseline, with reductions of 71–79% across four LoRA-adapted backbones. On the second benchmark, LoRA CADET achieves 46.7% leaf accuracy, 9.7 percentage points above the majority baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.