Before LLM Measurements Can Disagree: Auditing Output Admissibility Across Systems
Abstract
Large language models (LLMs) now serve as measurement instruments in socialscience and ML pipelines, and their outputs are compared across models and runs to assess reproducibility. Such comparisons presuppose an admissible output from every call. Yet recent work resamples failed calls or drops them before aggregation, so whether each system produced admissible outputs on the same inputs goes untested. We therefore audit output admissibility before agreement with a locally frozen, stop-rule-governed protocol. It treats every failed call as terminal and keeps verified zero extraction, abstention and terminal failure as separate states. The pipeline first extracts atomic stakeholder claims with verbatim evidence and then classifies them into seven frames. Its inputs are 11 masked documents from 8 carbon capture, utilization and storage (CCUS) project clusters. Three open-weight 7–8B checkpoints were each run in three seeded repeats with fixed prompts and decoding settings. The primary estimand contrasts crosssystem with within-system Jensen–Shannon divergence of cluster-level frame distributions. The frozen protocol requires every cell of every cluster to be complete and stops at the first terminal failure. The run stopped at its first call. We then executed all remaining calls to characterize the failure flow. Only 23 of 153 extraction calls returned schema-valid output with every claim anchored to verbatim evidence. Common support, all 9 cells complete, held in 0 of 8 clusters, so the primary estimand is not estimable. Most failures were evidence misalignments, and even dropping every non-conforming claim let only 2 of 8 clusters pass extraction. Failures recurred across independently seeded repeats even as responses varied, and their kind differed across systems. Token-level replay traced every unparseable constrained response short of the output-length cap to the constraint adapter. Recoding failures as empty extractions or abstentions gave 7 of 8 clusters apparent common support. The study concerns measurement reproducibility for this finite corpus and these checkpoints. We did not validate outputs against human labels, and the study does not assess semantic accuracy, construct validity or which model is correct. For these checkpoints, prompts and decoder, which inputs yielded an admissible output depended on the system. Admissibility should therefore be audited before cross-system agreement is reported.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.