acceptodds
Under review as a conference paper at ICLR 2027

Heuristic Bench: Evaluating LLM Agents by Authoring Executable Policies

Abstract

Large language model (LLM) agents have shown strong capabilities in agentic problem solving. But assessing these capabilities across models in a human-interpretable and fine-grained way remains challenging. Inspired by a recent paradigm, Heuristic Learning, we introduce Heuristic Bench, a benchmark in which agents iteratively solve tasks by authoring executable code as policies that can be run and evaluated independently. As code, these policies are human-interpretable and inspectable, allowing readers to examine the strategies behind the scores and investigate potential shortcuts. Using this benchmark, we evaluate 47 models with varying coverage across 45 problem-solving tasks, with discrete and continuous action spaces in solo, competitive, and cooperative settings. Controlled execution, independent evaluation, and a wide range of metrics support reliable measurement of model performance. The benchmark's breadth allows us to rank models by their ability, and its analyses reveal fine-grained traits, such as performance-cost trade-offs across models and gains from additional attempts, uneven gains across model versions, imperfect yet still informative self-reported scores, and partial agreement in model rankings across tasks. Interestingly, policy inspection reveals that capable models can exploit unintended shortcuts, including copying opponent and task code to construct exact-simulation search policies. These findings highlight the importance of careful evaluation design to reveal model capabilities and inform how AI models are evaluated, used, and developed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.