acceptodds
Under review as a conference paper at ICLR 2027

Output-Derived Supervision for Controllable Synthetic Text

Abstract

Attribute-conditioned text generators can produce fluent text that does not express the requested attributes. We ask whether a generator's own outputs can provide useful feedback for improving attribute control after access to source records ends. We generate text from requested attribute tuples and treat the requests as noisy labels for the resulting outputs. These pairs train a lightweight body-only verifier without new human or language-model annotations. The frozen verifier then guides reinforcement fine-tuning or best-of-eight candidate selection. On Enron, reinforcement fine-tuning improves automatic attribute recovery in all six seeds, with a mean gain of percentage points, and two independent language-model judges confirm the improvement on three evaluated seeds. An Amazon transfer study shows a -point independently judged gain. The improvements are attribute-dependent and reduce distributional fidelity. In a matched comparison, reinforcement fine-tuning achieves higher judged recovery than reranking, while reranking better preserves fidelity. These results show that requested attributes paired with generated outputs provide useful but imperfect control supervision without additional source access.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.