acceptodds
Under review as a conference paper at ICLR 2027

EviForge: Adaptive Visual Judge Models Evidence and Weakness Guided Training

Abstract

Generated-image evaluation must adapt its criteria to each prompt-image pair while remaining fast, affordable, and private enough for routine use. Distilling multimodal large language models into a local judge is practical, but teacher errors can contaminate supervision, while untargeted follow-up training may miss the student's actual weaknesses. We therefore introduce EviForge, a training method for lightweight, adaptive judge models with three components. Sample-adaptive facet judgment selects the applicable items from 60 facets under seven dimensions and returns scores, evidence-grounded rationales, and defect locations. Evidence-adjudicated multi-teacher distillation combines generalist judgments with routed specialist evidence: agreed facets are merged, while disputed claims undergo challenge and defense before separate adjudication. Weakness-driven refinement diagnoses recurring errors by facet, scene, error type, and direction and constructs targeted prompts and images. Verified examples then train the student by correcting individual facet errors, calibrating the complete judgment, and replaying historical data to retain learned abilities. We instantiate EviForge on Qwen3-VL-8B and Qwen3.8-27B. Across eight public datasets and blinded human evaluations, we assess facet selection, score consistency, evidence-grounded rationales, defect localization, image comparison and ranking, and generator improvement when the judge is used as a reward model. Both final models improve substantially over their corresponding backbones on most measures. For example, EviForge-8B improves defect-localization coverage by 63 percentage points and evidence hit by 31 percentage points. On most reported measures, both match or surpass the much larger Qwen3.8-Max and Kimi K2.6.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.