acceptodds
Under review as a conference paper at ICLR 2027

Check, Don't Trust: Evidence Contracts for Small Language Agents in GPU-Cluster Operations

Abstract

Operators of shared GPU clusters often did not write the scheduler or the jobs they are diagnosing, and a small, locally hosted language model is an attractive interface for their questions. But an ungrounded model can state numbers no metric supports, and an agent allowed to act can shut down a node that is still working. We study ClusterMind, an assistant built on 1–1.5B models in which a single retrieval pass produces one frozen, typed evidence set that is shown verbatim to the model, the operator and a deterministic checker, and in which state-changing actions pass a busy-node precondition re-evaluated at confirmation time. To evaluate it without synthetic metrics we build PAI-Diag from the Alibaba PAI GPU trace: 320 questions over 40 real cluster snapshots with computed gold answers, a 600-item controlled answer set, and trace-driven action scenarios. The prototype’s original grounding check certifies most wrong answers. A claim-level checker removes fabricated numbers but still delivers answers that are true of the evidence and wrong for the question; adding superlative, set and question-target claims raises the accuracy of delivered answers on held-out Llama-3.2-1B outputs from 38.5% to 98.5% at 24.5% coverage, with no model call and about 65 μs per answer. A 7B LLM judge given the same evidence is a weaker gate: at matched coverage its delivered answers are 75.4% correct, and it costs 10.7 s per answer on CPU. On actions, 47.0% of busy machines in the trace show ≤10% GPU utilisation, so a utilisation-based guard would permit shutting down nearly half of them, and 6.3% of idle candidates become busy within 30 minutes. The guarantees are narrow and we report where they fail, most notably when retrieval omits the relevant node

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.