acceptodds
Under review as a conference paper at ICLR 2027

Right Values, Wrong Records: Benchmarking and Improving Record Formation in Prompt-to-Data Systems

Abstract

Prompt-to-data systems can recover correct values but assign them to wrong records without shared source identifiers. End-to-end evaluation combines this decision with extraction and structural generation. RecordBench isolates record formation using controlled source-table candidates and document extractions: 2,382 episodes from 807 tables, including explicit-identity controls. Reference-Guided Assembly (RGA) uses a frozen tabular foundation model to infer candidate compatibility from reference records in context, then forms records by constrained global assignment, preserving observed values under complete or partial matching. Component-macro target-cell accuracy reaches 76.81% on controlled and 68.94% on extracted candidates; controlled RGA exceeds Qwen3.5-27B's executed tools by 18.23 percentage points. Same-backbone scoring ablations and separate assignment interventions establish the roles of reference-conditioned predictive distributions, cross-target conditioning and global competition in RGA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.