acceptodds
Under review as a conference paper at ICLR 2027

EAGLE: Grounding Multimodal LLM Judges for AI-Generated Image Assessment via Evidence Anchors

Abstract

Multimodal large language models (MLLMs) provide a promising foundation for AI-generated image assessment, yet existing MLLM-based evaluators typically derive assessment scores from holistic multimodal representations. This overlooks a fundamental distinction: different judgments require different evidence dependencies, e.g., text-image semantic correspondence depends on cross-modal evidence, whereas perceptual quality or authenticity relies primarily on visual evidence. To adaptively characterize these distinct evidence dependencies across different assessment tasks, we introduce evidence anchors that explicitly ground different assessment judgments in their relevant evidence. Building on this idea, we propose EAGLE (Evidence-Anchored Grounding for Multimodal LLM Evaluation), an MLLM judging framework that couples holistic MLLM scoring with anchor-guided evidence decomposition. EAGLE learns a set of latent queries divided into shared and anchor-specific groups: shared queries capture evidence reusable across judgments, while anchor-specific queries adaptively extract visual or cross-modal evidence to ground their corresponding judgments. This design preserves the broad assessment capability of pretrained MLLMs while providing each judgment with evidence tailored to its underlying dependency. Extensive experiments on popular benchmarks show that EAGLE achieves substantially improved within-dataset and cross-dataset generalization, validating the importance of explicitly grounding MLLM judges in their underlying evidence dependencies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.