acceptodds
Under review as a conference paper at ICLR 2027

RubricFly: Rubric-Guided Adaptive Evaluation and Judge Learning for Multilingual Translation

Abstract

As large language models improve at translation, static benchmarks lose diagnostic value: overall scores separate models less and reveal little about which capabilities remain weak. Sustaining that value requires tests that target specific capabilities, such as meaning preservation, register, and cultural adaptation, with question-specific criteria, and that are renewed as models improve. Yet translation benchmarks are mostly authored and updated by hand, and automatic generation for translation either optimizes difficulty alone or yields a fixed pool of items that is not renewed from observed model behavior. Model failures could guide renewal, but only if the Judge reliably separates genuine failures from valid translations; rejecting a valid rendering misstates a model's weaknesses and misdirects the next round of tests. We present Rubricfly, the first LLM-driven translation evaluation framework that combines automated question and rubric generation with feedback-guided benchmark renewal and human-guided Judge rule learning. Shared testpoints specify the capability under test and the response requirements that guide both construction and evaluation. A production loop turns model failures and coverage gaps into new questions and rubrics, and a Judge loop distills human corrections into reusable decision rules without updating model parameters. We build fly-corpus, a 500-question benchmark covering bidirectional translation and cross-lingual rewriting between Chinese and English, French, Japanese, Spanish, and Korean. Under a common GPT-5.5 Medium judging protocol, the 47 questions introduced by renewal average 43.32 across ten models, against 80.95 for the saturated questions they replace; full-feedback replacements are harder than those generated without model feedback (43.06 vs 51.25) and separate models more than difficulty-only ones. Rule-guided rechecking reduces essential-item errors from 307 to 291 under three-draw voting, more than rechecking without rules, although response-level errors are unchanged. Together, these results show how capability diagnosis can drive test renewal and how human corrections can improve the judgments on which renewal depends.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.