acceptodds
Under review as a conference paper at ICLR 2027

Continual Experience Learning from Expert Feedback for NL-to-Cypher Data Synthesis

Abstract

Natural-language-to-structured-query systems rely on paired natural-language requests and executable queries for training, prompting, and evaluation. Constructing such data manually is expensive, while LLM-based synthesis can introduce semantic errors that are not detected by syntax or execution alone. Moreover, corrections to individual examples are often discarded rather than reused to improve future synthesis. We study this problem primarily in the setting of natural-language-to-Cypher (NL-to-Cypher), where existing benchmarks remain relatively limited in semantic quality and structural coverage. We present CIRA, a continual experience-learning framework that converts expert corrections into reusable, scoped guidance for subsequent data synthesis. Instead of updating model parameters, CIRA compiles feedback into experiences that capture the failure, corrective guidance, and applicability conditions, then consolidates and selectively retrieves them during future generation. On an unseen CypherBench graph, CIRA generates NL-to-Cypher data with an estimated 95.0% validity based on a manually verified sample. Compared with 1,000 template-generated benchmark examples, CIRA produces 680 versus 503 distinct query structures and reduces structural duplication from 50% to 32%. In downstream fine-tuning, combining CIRA-generated and benchmark data yields the highest overall execution accuracy for both evaluated models. Across five rounds of expert feedback, CIRA increases the valid-pair rate from 0.29 to 0.92, while the same generation pipeline without experience memory ends at 0.21. On identical specifications, retrieving raw feedback reaches 0.67, showing that compiling and consolidating feedback provides additional benefit beyond retrieval alone. We further show that the learned experience bank transfers across a new schema, a different generator, and a new domain, improving the valid-pair rate by 0.21–0.55 over the corresponding memory-free pipeline. Finally, transferring eligible experiences from Cypher to SQL improves validity from 0.58 to 0.79, providing evidence that the learned guidance can extend beyond the query language on which it was acquired. The code is available at https://anonymous.4open.science/r/CIRA-3814/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.