Heuristic Bench: Evaluating LLM Agents by Authoring Executable Policies
Abstract
Large language model (LLM) agents have shown strong capabilities in agentic problem solving. But assessing these capabilities across models in a human-interpretable and fine-grained way remains challenging. Inspired by a recent paradigm, Heuristic Learning, we introduce Heuristic Bench, a benchmark in which agents iteratively solve tasks by authoring executable code as policies that can be run and evaluated independently. As code, these policies are human-interpretable and inspectable, allowing readers to examine the strategies behind the scores and investigate potential shortcuts. Using this benchmark, we evaluate 47 models with varying coverage across 45 problem-solving tasks, with discrete and continuous action spaces in solo, competitive, and cooperative settings. Controlled execution, independent evaluation, and a wide range of metrics support reliable measurement of model performance. The benchmark's breadth allows us to rank models by their ability, and its analyses reveal fine-grained traits, such as performance-cost trade-offs across models and gains from additional attempts, uneven gains across model versions, imperfect yet still informative self-reported scores, and partial agreement in model rankings across tasks. Interestingly, policy inspection reveals that capable models can exploit unintended shortcuts, including copying opponent and task code to construct exact-simulation search policies. These findings highlight the importance of careful evaluation design to reveal model capabilities and inform how AI models are evaluated, used, and developed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.