When the Harness Confirms the Hypothesis: Silent Configuration Faults in Multi-Agent LLM Evaluation
Abstract
Reported gains and losses from multi-agent LLM pipelines are only as trustworthy as the harness that produced them. We audit one multi-agent evaluation end to end and find seven faults that produced no error and no crash, but changed conclusions: a chat template that silently discarded every system prompt, leaving two condi- tions differing only in that prompt byte-identical on 721 of 721 problems; a gram- mar argument swallowed by a constructor, making the “grammar-constrained” condition a byte-identical duplicate of the unconstrained one; a decoding tempera- ture that never reached the sampler, so a run recorded as greedy sampled at 0.2–0.6 against genuinely greedy baselines; hidden per-role token caps truncating one role on 77–90% of problems; and a grader deviating from the benchmark’s standard protocol, worth up to 11.6 points and biased toward one condition. Every fault produced plausible results that confirmed the hypothesis under test, which is why each survived days of monitoring. We characterise the faults, quantify each one’s effect by re-running the affected cells, and release detectors that find them from outside any framework via a recording proxy. We then report corrected results for the pipeline under study on two 7–8B models and three benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.