Beyond the Order Quantity: Benchmarking Behavioral Biases of LLM Inventory Agents under Uncertainty
Abstract
Large language models (LLMs) can solve quantitative problems, but inventory agents must act under uncertain demand, lead times, and partner behavior repeatedly. Existing evaluations, however, rarely distinguish limited task competence from systematic behavioral bias or move beyond a single inventory setting. We introduce AIM-Bench, a benchmark that evaluates LLM inventory agents across five environments, from a single-period newsvendor problem to multi-period and multi-echelon supply chains. AIM-Bench varies the source and structure of uncertainty and evaluates agents at three levels: operational outcomes, distance from an ex-post optimal action, and behavioral signatures such as mean anchoring, demand chasing, and the bullwhip effect. Experiments with eight state-of-the-art LLMs reveal that strong reasoning ability does not ensure reliable replenishment, because even the strongest models retain systematic behavioral biases: some models still anchor orders toward mean demand, and all evaluated models amplify demand variability in the Beer Game. These biases can remain hidden when models are compared only by aggregate cost or stockout rate. We further test two lightweight interventions. Cognitive-reflection prompting reduces anchoring for several models, information sharing substantially weakens the bullwhip effect of some models. None of these remedies is uniformly effective. Our results show that LLM inventory agents should be evaluated not only by what outcomes they achieve, but also by how their decisions respond to uncertainty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.