DynaPlanBench: Benchmarking Long-Horizon Agentic Planning in Dynamic Environments
Abstract
LLM-based agents are increasingly expected to solve long-horizon tasks through interaction with dynamic real-world environments. However, existing benchmarks largely evaluate plan generation or task completion in static settings, or introduce dynamics only as isolated state changes and disruptions. Such settings do not fully capture a defining property of real-world execution: the environment evolves over time, and previously executed actions may create persistent commitments that continue to constrain future decisions. To address this gap, we introduce DynPlanBench, an executable benchmark for long-horizon agentic planning and execution in dynamic environments. Across 153 tasks spanning logistics, retail, and travel, DynPlanBench models two complementary dynamics: Time-Indexed Dynamics, where environment states and action outcomes vary with execution time, and Event-Driven Dynamics, where task-specific disruptions arise during execution. Together, these dynamics create a continuous planning-and-execution process, requiring agents to perform temporal planning over world states at future execution times and commitment-aware replanning under persistent commitments from prior execution. Evaluation of 20 frontier LLMs shows that even the best-performing model achieves only a 48.7% success rate. Further analysis reveals two recurring failure modes: global temporal inconsistency across dependent steps and commitment-state drift across successive repairs. These findings identify temporal consistency across evolving world states and commitment-state maintenance across successive plan repairs as key challenges for reliable long-horizon agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.