The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
Abstract
LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on local disk, including logs, context snapshots, checkpoints, and debug traces. These bytes are absent from the leaderboards we survey, yet bear on whether an agent system can be deployed on desktops, under data-retention compliance regimes, or at fleet scale. To our knowledge, AgentFootprint is the first systematic cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite covers total retention, channel composition, duplication, growth exponent, compressibility, and a conversation-history reconstructability score. It addresses a measurement trap: naïve byte-level measurement understates duplication by an order of magnitude, because database paging and JSON escaping obscure repeated content. The footprint has two determinants: the logical volume generated by the agent's behavior and the amplification added by the persistence layer. A fixed-trace control compares the latter: the same scripted content replayed through each persisting framework's adapter yields retained sizes differing by 6.7×. Among persisting defaults achieving 100% accuracy under identical models, tools, and tasks, retained bytes differ by 15.7×. The defaults bundle different recovery and audit capabilities, so the suite reports storage jointly with accuracy and reconstructability. Three full-history configurations grow superlinearly on a repeated-observation stress task, and on a deliberately minimal write task framework residue exceeds the delivered output files by orders of magnitude. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude in per-instance volume with no detectable correlation with resolve rate. A content-addressed store reduces retention 4.8–32.7× while preserving all properties checked by our restoration tests, including every reconstructability score. Reconstructing recorded conversation content does not require megabytes, and these storage metrics can be reported alongside inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.