acceptodds
Under review as a conference paper at ICLR 2027

SandboxLive: Benchmarking LLM Agents on End-to-End Dynamic Malware Analysis

Abstract

LLM agents are increasingly used as security analysts, yet it remains unclear whether they can analyze an unseen malware sample end to end, i.e., triage it, plan and run its detonation in a sandbox, and report only what the run supports. Existing benchmarks evaluate agents either inside environments that are prepared for them or on static and question-answering tasks, leaving untested whether an agent can run the detonation of a real sample itself and ground its report in what the run observed. In this paper, we present SandboxLive, a live benchmark and an open harness for dynamic malware analysis by LLM agents. Each season draws Windows and Linux samples first seen after the release of every evaluated model, and the evaluated model plays four agent roles that operate a real sandbox under a deterministic controller. After the experiments, each analysis report is scored against the evidence of its own run by deterministic verification rules, without an LLM judge. Over three seasons (119 samples, eight models, and 675 runs), our results reveal three key findings. (1) The leading models are close, and they differ in how they interpret the evidence rather than in how they set up the detonation. (2) Only 26 of the 63 failed runs are attributable to the model. (3) How the harness enforces a constraint, rather than how often a model violates it, determines the cost of a violation. Code and season lists are provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.