STED: MEASURING CONSISTENCY IN LLM STRUCTURED OUTPUTS
Abstract
Structured responses are common in industrial applications of large language models (LLMs), with tool-calling agents a prominent use case. Yet even at temperature , repeated prompts can produce inconsistent tool choices and argument values despite parseable responses. We propose STED (Semantic Tree Edit Distance), a type-aware score for comparing hierarchical structured outputs. It matches terminal values at corresponding tree paths and aggregates their graded agreement using soft precision and recall. On 208 structural human comparisons used during development, STED makes 88.0% correct decisions versus 79.8% for a field-path similarity baseline, counting metric ties as unsuccessful decisions. Frozen numerical-equivalence tests isolate a benefit of interpreting numerical literals, with within-value-group AUCs of .999–1.000 versus .612–.619 for a control with matched text processing. Textual relatedness and external model judgments show no consistent improvement. We apply the fixed score to 2.1M repeated generations from 18 LLMs, measuring consistency separately from validity and prompt coverage. The observed rankings and model-dependent schema-enforcement effects show why repeatability adds information to model selection alongside task quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.