Large Language Model-Based Game Agents Adjust Their Play to Graded Instructions for Competing Priorities
Abstract
Automated playtesting may require agents to emulate human players at different skill levels, whereas game agents used as non-player characters may need to exhibit a consistent persona. These uses may call for different balances between playing strength and other desired behaviors. For example, in chess and shogi, preserving a win and favoring piece captures need not lead to the same move. We therefore investigate whether and how large language model (LLM)-based game agents respond to graded natural-language instructions specifying the relative importance of these goals. We test game agents built on three LLMs in Animal Shogi, Breakthrough, and Rush Hour on states where following the requested style costs playing strength. We vary relative priority over five levels using numerical, modal, and adverbial instructions. Across all tested models and games, agents exhibit a graded behavioral response, with style-matching actions becoming more frequent overall as the requested style is assigned higher priority.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.