acceptodds
Under review as a conference paper at ICLR 2027

XGenerator: Skill-Guided Synthesis of Verifiable Training Data for a Low-Resource Modeling Language

Abstract

Low-resource domain-specific languages (DSLs) often lack training data linking explicit requirements, executable models, and verification evidence. In industrial system modeling, constructing such data requires compilation, simulation, and requirement checking within a unified workflow. We introduce XGenerator, a skill-guided framework for synthesizing verifiable training data that keeps base-model parameters fixed during generation. Versioned procedural skills organize language specifications, modeling knowledge, and verified experience for on-demand retrieval. The language model produces a compact executable intermediate design describing components, states, and event interactions. Deterministic tools translate it into source code and perform compilation, simulation, and requirement checking against frozen behavioral contracts. Failure evidence guides scoped design repair and revalidation while requirements and shared skills remain fixed. The resulting artifacts link requirements, designs, code, and execution evidence; subsequent skill revisions require separate development validation. We evaluate XGenerator on 80 discrete-event modeling tasks across eight domains and five difficulty levels in X-language, a low-resource DSL. XGenerator improves first-pass requirement satisfaction over prompt-only baselines by 88.75 and 40.00 percentage points on GPT-6-astra and DeepSeek-flash, respectively. With at most one feedback-guided repair, the final requirement satisfaction rate reaches 97.5% and 80.0%. We further construct a model dataset for fine-tuning. Across four Qwen3-8B runs, requirement satisfaction on 410 tasks with executable criteria reaches 68.78–70.24%, compared with 0% for the base model under raw-output evaluation. These results support constructing traceable training data for low-resource modeling DSLs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.