Benchmarking Tool-Dependent Safety in Agentic Vision-Language Models
Abstract
Vision-language models (VLMs) are shifting from standalone models to agentic VLMs that autonomously use external tools to solve tasks that their built-in knowledge and capabilities alone cannot handle. As this expands what models can solve, it should also expand what risks they can recognize. Current VLM safety benchmarks, however, rarely test this capability because the safety implications of their samples are usually apparent from the original input. We introduce SAVIT, a benchmark where an appropriate response depends on safety-critical information that is difficult to obtain without effective tool use. SAVIT separately evaluates (1) whether models obtain this information and (2) whether they respond appropriately based on it. We construct manually reviewed samples across safety categories and evaluate recent VLMs, showing that models pass both evaluations on only of samples on average. Our analysis of tool-use behavior and failure modes reveals that using tools well matters more than using them often, and that failures arise throughout the agentic process, from recognizing when tools are needed to using the obtained information in the final response.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.