NetMend: Can LLM Agents Repair Faults in Live Networks?
Abstract
As LLM agents move into infrastructure operations, rigorous evaluation of their network-troubleshooting ability becomes critical. However, most existing benchmarks evaluate diagnosis rather than repair. Others evaluate repair, but in simplified settings far removed from operational networks. We present NetMend, a benchmark that evaluates whether LLM agents can repair injected faults in live containerized networks. The agent starts from a symptom report with the injected fault hidden, and must localize and repair it using tools that run real operator commands. Correctness is judged from the state of the running network, rather than from the agent's own report. NetMend also scores whether the agent found the true root cause and whether its fix damaged unrelated parts of the network. Underlying NetMend is a task generator rather than a fixed dataset. It synthesizes networks across blueprints, routing protocols, and network operating systems. Each task is emitted as a replayable specification with a computed difficulty and a minimal reference solution. Beyond synthetic networks, NetMend also rebuilds two production topologies from real vendor configurations. NetMend currently comprises 29 fault classes and their compound combinations, instantiated on networks of configurable scale. Across five open-weight models, repair success ranges from 25.3% to 85.1%, and even successful agents take four to eight times as many tool calls as the minimal reference solution. Nearly one in five correctly diagnosed faults is still not repaired, while no repair succeeds without the correct root cause. Code is available at https://anonymous.4open.science/r/NetMend-0517.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.