Snapshot-correct, replay-wrong: Grading generated data pipelines over time
Abstract
Agent-written data pipelines can pass the test they ship with and still corrupt the tables they maintain: executed a second time on the same records, or on records delivered late, twice or out of order, they go silently wrong, and a benchmark that runs a generated artifact once cannot see it. Grading eleven model arms on 40 tasks we built against a fixed snapshot, as seven of nine agent benchmarks we audit do, passes 86.0–100.0% of the pipelines that run; re-executing those same artifacts under duplicate, delayed, reordered and retried delivery finds 7.0–79.2% of the passing ones wrong (1.7–78.1% without five tasks an independent reviewer found ambiguously worded). dbt's default grain tests, placed on each task's key, catch 4 of 110 such failures: a uniqueness test on a business key is satisfied by exactly the upsert that double-counts. In exploratory analyses, the algebra of the update predicts which hazard class a pipeline fails, though not how often: additive folds double-count a redelivered record where naturally idempotent ones absorb it, a gap of +42.2 [+26.6, +58.4] points that survives a port to Python. The defect also has a locus we can move: rewriting uniqueness to the event-arrival boundary raises idempotency pass rates 74.7 and 67.0 points on DuckDB and PostgreSQL where an entity-keyed rewrite and a syntax-preserving one do not. Telling a model about the property narrows the gap without closing the spread between models: the share of passing pipelines that are wrong falls in all 11 arms (significantly in 10 after adjusting for multiplicity) and reaches zero on 3, yet all-in (silent-or-crash) failure still spans 0.0–71.0%, and on Sonnet 5 much of the fall is silent wrongness traded for crashes on a constraint the model itself declares. Evaluating stateful generated code therefore means perturbing its execution history rather than grading one snapshot of it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.