Writer-R1: Dynamic Rubric Mining for Reinforcement Learning of Open-Ended Writing
Abstract
Training models for open-ended writing requires diverse prompts and sufficiently precise specifications of response quality. Constructing such specifications at scale entails recovering task requirements, consolidating overlapping criteria, and defining observable differences between performance levels. We introduce Writer-R1, a writing training pipeline built around Inductive Rubric Generation (IRG), which integrates these operations into an automated rubric construction procedure. The pipeline expands an open writing corpus and uses the resulting prompt-specific rubrics throughout supervised training and reinforcement learning (RL). Each rubric is concatenated with the user request and also defines the reward criteria, making the requirements used for assessment available to the writer before generation. A compact model, RubricGen-4B, learns to supply rubrics for new requests. Writer-R1 improves WritingBench scores by 0.182 and 0.158 over the 4B and 8B backbones, respectively, with gains also on HelloBench. Training ablations favor IRG over direct rubric generation, and further analyses support rubric transfer and the robustness of writer improvements across evaluators. These results support using inductively mined rubrics as both generation guidance and reward specifications for open-ended writing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.