FROM DATA LAKES TO TRAINING TASKS: GROUNDED AND CONTROLLABLE SYNTHESIS FOR DATA AGENTS
Abstract
Training data agents at scale rests on a steady supply of grounded, verifiable data-analysis tasks. Real-world data lakes make these hard to produce: they hold far more files than can be inspected individually, and the few files that share a meaningful join path cannot be identified from local context alone. We propose DATA2TASK, which represents a data pool as a compact physical-semantic dual-view graph: its physical view replaces exhaustive inspection with selective access, while its semantic view aligns columns denoting the same concept across files and so surfaces latent relations. DATA2TASK then samples a provenance-preserving subgraph and instantiates a compatible analytical operator over it, forming a task motif whose answer is computable from the sampled files and therefore grounded by construction; a staged pipeline turns each motif into a natural-language request with executable ground truth, kept only when independent solvers reproduce that answer, and scored separately for structural and semantic difficulty. With only 2,000 verified tasks and selected solution trajectories, fine-tuning Qwen3-8B achieves an average score of 55.17% across four held-out data-analysis benchmarks, outperforming the 29x larger Qwen3-235B and matching the 60x larger Qwen3-Coder-480B on TableBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.