acceptodds
Under review as a conference paper at ICLR 2027

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Abstract

Language-model agents increasingly act as persistent coworkers that work alongside users across many days, during which the surrounding environment keeps changing on its own and key evidence arrives in many modalities. Existing benchmarks miss this regime, scoring agents within a single static, largely text-centric episode. We introduce ClawMark, a benchmark for coworker agents built around multi-turn, multi-day tasks whose sandboxed service state evolves between turns, with fully rule-based verification and no LLM-as-judge. ClawMark comprises 100 tasks across 13 professional scenarios, scored by 1,537 deterministic checkers over post-execution service state. Benchmarking seven frontier systems shows the setting is far from solved: the strongest reaches 75.8 weighted score but only 20.0% strict Task Success, and performance drops sharply after the first exogenous environment change. ClawMark thus isolates adaptation to changing state as a central open challenge; we release the benchmark, evaluation harness, and construction pipeline to support reproducible coworker-agent evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.