acceptodds
Under review as a conference paper at ICLR 2027

Is LLM-Curated Policy Evidence Trustworthy Enough for Causal Analysis? An Audit-then-Propagate Framework Applied to Global Digital-Nomad Visa Data

Abstract

Large language models (LLMs) can turn legal and policy documents into structured datasets, but extraction accuracy alone does not show whether the resulting evidence supports a downstream causal conclusion. We introduce an **audit-then-propagate** framework that independently checks policy claims against primary sources and measures how a field-specific error affects staggered-adoption difference-in-differences estimates. In a digital-nomad-visa study spanning 190 jurisdictions, a human audit confirms 33 adopters. Among the LLM pipeline’s 46 determinate program-existence calls, 32 agree with the audit (69.6%; Cohen’s ); 144 jurisdictions remain unresolved by the pipeline’s evidentiary standard. Adoption year matches in five of six fully cross-checked cases, while tax and financial-threshold fields lack independently verified reference values. A 200-iteration semi-synthetic experiment substitutes recorded announcement dates for verified effective dates. This timing error attenuates estimates toward zero by 0.25–0.28 for both two-way fixed effects and a Callaway–Sant’Anna-style estimator, reducing 95% confidence-interval coverage from 72.0% to 50.0% and from 90.0% to 67.5%, respectively. An application using verified dates and national unemployment data produces directionally negative event-time estimates, but none is individually distinguishable from zero; the specification also provides no pre-treatment estimates with which to assess differential trends. The study shows how an auditable curation error can materially change inference, while making the benchmark’s unresolved coverage and the limits of its observational application explicit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.