PAVES: Prior-Guided Dataset Compilation for Effective Training and Scalable Expansion
Abstract
High-quality tabular datasets are often too small for effective model training because representative records are costly to collect. Synthetic expansion can increase the training size, but finite observations may contain spurious cross-column relationships. Sampling more records from a fixed generated distribution reduces variance without removing this relation bias, causing downstream models to learn misspecified structure more consistently. We introduce PAVES (Prior-guided Acquisition and Validation of Executable Specifications), a framework for prior-guided relation-space compression and executable dataset generation. Prior knowledge narrows the candidate relationships, while observed data determine which candidates are retained and estimate their parameters. Under coverage and separation conditions, compression reduces the uniform finite-sample error of relation validation; we further connect the remaining relation error to task-relevant distribution shift. PAVES compiles validated relationships and statistically estimated marginals into an editable executable specification that can be reused to replay a dataset at its source size or generate larger training sets without retraining an implicit generator. Across tabular benchmarks, PAVES achieves strong statistical and specification fidelity and downstream utility, while target-size experiments show continued gains when the same specification is reused at larger scales. Code is provided in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.