acceptodds
Under review as a conference paper at ICLR 2027

DataStamp: Domain-Specific Data Synthesis via Structural Projection

Abstract

Large language models (LLMs) achieve strong performance on general-purpose tasks through instruction tuning, yet high-quality instruction–response pairs remain scarce in specialized domains such as humanities and medicine, where typically only raw, unlabeled documents are available. The central challenge is synthesizing pairs that are both grounded in target documents and diverse in reasoning patterns. Existing approaches either bootstrap from seed examples via LLM generation, defaulting to shallow and repetitive questions, or rely on manually engineered templates that offer limited diversity and fail to transfer across domains. We observe that high-quality QA data spanning diverse domains has already been produced by prior research and that each pair implicitly encodes a domain-free reasoning pattern. Inspired by meta-learning, we propose DataStamp, which treats these patterns as transferable prototypes and adapts them to new target domains. Each document is abstracted into a logic graph that retains only its logical skeleton; from existing cross-domain QA data, DataStamp extracts transfer graphs—typed-graph prototypes encoding how a logical structure should be transformed into an instruction–response pair. Given a target document, the most structurally compatible transfer graphs are retrieved via a relational GCN and projected onto the target content, producing grounded and diverse pairs without domain-specific engineering. Evaluated on twelve benchmarks across two domains, DataStamp outperforms open instruction corpora and domain-specific baselines, matching or exceeding Qwen3-8B-Instruct trained on proprietary data, while producing data that is substantially harder and more diverse than existing alternatives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.