Same Prices, Different Responses: Counterfactual Response Surfaces in Learning Markets
Abstract
Learning sellers can produce similar price paths for very different reasons. Some may act independently, while others may react strongly to a rival's action. Observed market outcomes alone may therefore hide important differences in strategic behavior. We study this problem by measuring how a seller responds when a rival's action is changed. For each unilateral intervention, we compare an intervention rollout with a matched control rollout that starts from the same market state and uses the same future shocks. Repeating this comparison across intervention settings gives a counterfactual response surface that describes the strength and duration of the response. We show that observed trajectories alone cannot identify these responses when different policy systems generate the same observed outcomes. Matched intervention and control rollouts make the responses identifiable and allow them to be estimated over a finite intervention grid. Experiments with scripted and learned pricing policies show that the response surface separates policy systems that look similar from observed prices alone. It also improves identification under the same evaluation budget, predicts responses to held-out interventions, and improves targeting when intervention capacity is limited. These results show that how a policy reacts to a controlled change can reveal information that its observed market outcomes do not.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.