MuseFACET: Benchmarking Trustworthy and Controllable Image-to-Text Generation for Musuem Collection Interpretation
Abstract
Museum collection interpretation represents a challenging task for Multimodal Large Language Models (MLLMs), requiring accurate object identification and knowledge-grounded, context-aware generation beyond literal visual descriptions. However, existing efforts on open-ended text generation for artworks and museum objects remain limited in dataset scale and diversity or in evaluation protocols that rely largely on reference-based similarity metrics. To address these limitations, we introduce MuseFACET, a benchmark comprising 5,000 museum objects from 9 museums. MuseFACET presents two hierarchical tasks: Identification task for predicting object attributes from images and Interpretation task for generating coherent, context-aware descriptions. We further propose an evaluation framework emphasizing trustworthiness and controllability through identification accuracy and five interpretation criteria: Faithfulness, Completeness, Conciseness, Depth, and Acuity. Evaluation of 14 proprietary and open-source MLLMs shows that the highest accuracy on the Identification task reaches only 31.2%. For Interpretation, models exhibit persistent limitations in factual grounding and object-specific insight under the image-only setting, while providing metadata substantially improves performance. Model performance also varies across cultural regions, revealing uneven capabilities in cultural interpretation. MuseFACET provides a systematic benchmark for museum collection interpretation and a concrete evaluation protocol for trustworthy and controllable open-ended image-to-text generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.