Rich-Node Graphs for Corpus-Grounded Long-Form SFT Data Synthesis
Abstract
Domain-specific supervised fine-tuning (SFT) requires training data that cover related subtopics while remaining grounded in source evidence. Existing pipelines often rely on parametric knowledge, individual passages, or flat retrieved contexts, limiting their ability to organize distributed evidence into coherent supervision. We present RiNoGen, a framework that transforms domain corpora into mixed-length question–answer pairs using provenance-aware Rich-Node graphs. RiNoGen selects corpus-supported seed entities, expands them with retrieved evidence, and applies Map–Reduce synthesis to compose local subtopics into globally coherent answers. The target Qwen3-8B model generates the grounded answers and is then fine-tuned on them, without requiring a larger teacher for answer generation. On a corpus of 670k Chinese educational papers, RiNoGen constructs 30,000 SFT examples. Under a fixed 3,000-example budget, it achieves 87.59% closed-book accuracy on a 427-item evaluation subset, outperforming the strongest baseline by 1.14 percentage points. Ablations confirm the contributions of graph expansion and Map–Reduce synthesis. An optional web branch further improves responses to time-sensitive topics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.