The Crowd Already Votes: Learning Creative-Writing Rewards from Human Interaction
Abstract
Creative writing aims to produce automatically open-ended text that is coherent and engaging narratives. Despite their strong text generation capabilities, large language models (LLMs) still struggle to consistently produce writing that receives high evaluations from human audiences. A central obstacle is that writing quality is highly subjective and real-world judgments emerge from collective rather than individual preferences. Consequently, existing models poorly approximate reader responses. We introduce REALFEEDBACK, a data-construction paradigm that treats online communities as continuously operating sources of preference supervision. It converts naturally occurring human interactions into reward signals without new large-scale annotation. We instantiate REALFEEDBACK in two domains: (1) For story generation, inline critiques yield passage-level process labels. (2) For forum discussions, nested discussions yield outcome-level preference pairs. We use these data to train REALFEEDBACK-PRM and REALFEEDBACK-ORM. Across established benchmarks and our constructed test sets, REALFEEDBACK-REWARD achieves state-of-the-art performance and outperforms much larger models. We also adopt REALFEEDBACK-REWARD in downstream test-time scaling applications for best-of-n (BoN) selection and find that it always selects outputs that receive higher human preference scores. These results position online communities as sustainable infrastructure for training reward models that better reflect collective human judgment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.