ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
Abstract
Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static prompts and clean-state environments miss failures that arise only when agents operate over persistent, partially completed workflows. Existing interactive benchmarks have advanced agent evaluation significantly, but do not systematically test whether agents preserve valid artifacts, repair stale state, replace incorrect objects, or reconcile conflicting evidence before acting. We present ClawForge, a generator-backed benchmark framework for executable command-line workflows under state conflict. The framework compiles scenario templates, grounded slots, initialized state, reference trajectories, and validators into reproducible task specifications, then evaluates agents step by step using normalized end state and observable side effects rather than exact trajectory matching. We instantiate the framework as ClawForge-Bench, a 362-task snapshot spanning 17 scenarios and 6 ability categories. Results across 11 model endpoints show that strict full-pass accuracy ranges from 35.9% to 61.6%, Wrong-State Replacement ranges from 0.0% to 52.2%, and Interrupted Workflow Resume ranges from 16.7% to 100.0%, revealing strongly non-nested capability profiles. Partial-credit and step-efficiency analyses further show that many failures are near-miss closures rather than early breakdowns, and that models exhibit qualitatively different failure styles under state conflict.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.