VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
Abstract
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate image-valued evidence—crops, masks, overlays, zoomed regions, and analytic renderings—that must remain addressable without accumulating unboundedly in multimodal context.We introduce VLM-in-Sandbox, a training-free framework for agentic multimodal reasoning in controlled computer environments. Its Visual Workspace registers generated artifacts in an image ledger, maintains a bounded active visual context, and lets the model explicitly promote selected evidence for subsequent inspection. This separates visual evidence generation, performed by sandbox tools, from visual evidence management.Across seven benchmarks and four VLMs, VLM-in-Sandbox achieves higher sample-weighted accuracy than Vanilla VLM and Append-only Sandbox for every model. A compiler-matched study on 1,260 examples isolates visibility from retention: the full policy reaches 66.27% accuracy with 18.6% fewer total tokens than the automatic, retain-all control. Paired evaluation of all 6,350 GPT-4.1-mini examples yields 302 rescues and 142 regressions relative to Original Append-only. A local vLLM study with prefix caching also reduces uncached prompt tokens, time to first token, and end-to-end latency, supporting explicit visual evidence management for sandboxed VLM agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.