acceptodds
Under review as a conference paper at ICLR 2027

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Abstract

Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in this broader setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables evaluation of heartbeat-conditioned proactive assistance, requiring agents to infer what to recommend at a fixed check time. Experiments show that GPT-5.5 achieves only 34.5% Pass@1, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments; fine-tuning Qwen3.5-27B on successful trajectories improves Pass@1 by 23.7 percentage points, demonstrating the pipeline's utility for generating training data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.