Concordia: Self-Improving Synthetic Tables for Federated LLMs
Abstract
Adapting large language models (LLMs) to tabular prediction in regulated cross-silo settings is difficult when institutions cannot share raw records, validation sets, synthetic tables, generators, or downstream model updates. This challenge is amplified by non-IID client distributions, extreme class imbalance, and limited communication budgets, where only one or two rounds may be feasible. Synthetic tables offer a natural on-premise training substrate, but existing pipelines often treat generation as static or optimize proxy objectives disconnected from the evolving downstream learner. We propose Concordia, a tri-level limited-shot framework that turns private validation utility into an online signal for refining client-local synthetic generation. Each client performs local LoRA adaptation on synthetic tables, learns a capacity-limited utility scorer from private validation feedback to reweight synthetic samples, and uploads only this scorer. The server pools and redistributes scorers within each task, and each client uses the pooled scorer feedback to refine its own generator with group-relative policy optimization (GRPO). Across finance and healthcare tabular benchmarks, Concordia improves early-round MCC under severe heterogeneity and imbalance, with the clearest gains on the long-tail Travel task where positives are near 0.01; the finance tasks are single-client streams and measure local self-improvement, while healthcare evaluates multi-client scorer pooling. Generator audits further suggest that scorer-guided refinement changes the support of synthetic training data, preserving coherent feature dependencies in finance and emphasizing multi-feature risk patterns in healthcare.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.