Meta-Rubrics: Learning to Discover Rubrics at Deployment for Efficient Task Adaptation
Abstract
Large language model (LLM) agents are increasingly deployed in open-ended environments where task objectives are under-specified and no task-specific evaluator is available. Reinforcement learning (RL) and inference-time search have driven major gains in verifiable domains such as math and coding, but these gains rely on a verification–generation (VG) gap: checking a solution is easier than producing one. In open-ended tasks, however, this gap cannot be assumed, because a verifier must first infer what constitutes success. This raises a central question: can an agent can construct or widen such a gap at deployment time from a small amount of environment interaction and use this gap to aid policy adaptation? We formulate this as a reward discovery setting and propose a framework, Meta-Rubrics, which infers a reward function, represented as a natural-language rubric, which can accurately assign credit to guide policy behavior at deployment time. In this framework, a rubric generator infers task-specific evaluation criteria, and an LLM judge applies these criteria to produce scalar scores and criterion-level feedback for candidate trajectories. The generator and judge are jointly optimized with RL over branched rollouts so that induced scores recover oracle reward orderings on held-out trajectories. At deployment time, this inferred reward can rank candidate responses or supply a signal for lightweight policy optimization with RL. We evaluate Meta-Rubrics on user simulation in Human-LM and curating data for a language model for an unseen task with a fixed fine-tuning budget. In both settings, Meta-Rubrics outperforms the highest non-oracle adaptation approaches such as in-context learning, evolutionary search, and rubric-based fine-tuning baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.