acceptodds
Under review as a conference paper at ICLR 2027

VisualClaw: A Real-Time, Personalized Agent for the Physical World

Abstract

Vision language models (VLMs) are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cost when processing dense video frames and long prompts, the agent scaffold remains static after deployment, and standard video-QA benchmarks do not test whether agents can use visual evidence inside tool-using workspaces. We present VisualClaw, a self-evolving multimodal agent built around two principles. First, hybrid encoding reduces deployment cost by filtering less informative streaming frames with a cascaded gate and compressing the text skill bank through hot/cold top- injection. Second, skill evolution lets the agent learn from failures: retrieved memories condition an offline evolver either as direct concatenated context or as guided evidence, producing skill-bank updates that help future questions. Across video-QA benchmarks with VLMs (Gemini 3 Flash and GPT-5.2), VisualClaw reduces API cost by versus full-frame upload and by versus offline Uniform-8 under the same skill bank. Under the streaming settings, our VisualClaw improves over the baseline by on average, while reaches a peak boost on EgoSchema with Gemini 3 Flash. To address the benchmark gap, we further curate VisualClawArena, a -scenario multimodal agentic benchmark built through a strict five-stage pipeline; models must use video evidence, documents, dynamic updates, and executable checks inside a workspace. On VisualClawArena, the same self-evolution framework with computer-use agent backends improves macro accuracy by for Codex (GPT-5.5) and for Claude Code (Sonnet 4.6) over no-evolution baselines, with a cost reduction compared to the uniform-sampled baseline. These properties make VisualClaw a natural fit for streaming edge applications such as AI glasses, where the cascade filters continuous video locally and transmits only selected keyframes at interaction, while self-evolution supports adaptive assistance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.