CookBench: Can Digital Agents Actually Cook? A Multimodal Benchmark for Long-Horizon Embodied Planning
Abstract
Recent multimodal agents have demonstrated substantial progress in digital environments, yet whether these capabilities transfer to long-horizon, physically grounded tasks remains unclear. To bridge this evaluation gap, we introduce CookBench, a high-fidelity interactive benchmark for long-horizon embodied planning in complex cooking scenarios, in which agents act primarily through structured text observations, with RGB screenshots as an optional input channel. Featuring 133 single-dish tasks with an average length of over 130 steps, CookBench goes beyond existing benchmarks by incorporating a large action space, 176 interactive objects with 50 functional items, and high-fidelity physical dynamics. Through systematic evaluation of representative agent frameworks and state-of-the-art vision-language models via a standardized evaluation harness (the Embodied Planning Machine, EPM), we reveal a substantial reality gap: even GPT-5.4-driven agents reach a dish-quality score of only 8.35 out of 100, while human operators reach about 93 under the same interface, showing that near-perfect scores are attainable through it. Ultimately, CookBench provides a crucial testbed for this digital-to-physical transition, tracing execution breakdowns to three fundamental capability deficits: absent physical grounding, poor closed-loop feedback comprehension, and severe embodied state amnesia. Code and benchmark resources are available at https://anonymous.4open.science/r/cookbench-anon-C13C.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.