Visual Dependence and Utility in Multimodal Knowledge Graph Completion: A Controlled Pilot Study
Abstract
Multimodal knowledge graph completion experiments often use changes in aggregate ranking metrics to assess the value of visual information. Interpreting these changes requires specifying which inputs, model parameters, and evaluation labels remain fixed. We report a controlled pilot study that examines these choices through three diagnostics: alpha-channel image preprocessing, positive-label filtering with fixed scores, and visual-package reassignment in separately trained MyGO models. Correcting transparent-image handling changes hundreds of source images, while absolute mean new-target MRR changes remain below in our fixed linear-transfer setup under its original validation filter. Adding withheld training positives to the evaluation filter changes MRR without changing model scores and reverses one descriptive fusion comparison. In a separate transductive experiment, five new visual mappings reduce the original-image model's MRR by 0.0544–0.0587. The permanently permuted model changes by less than 0.0006 in either direction, despite more than 58% of query ranks changing under every mapping. A zero-visual control remains exactly invariant. These observations distinguish input dependence, net ranking utility, and evaluation sensitivity within the tested configurations. The experiments are exploratory: the MyGO comparison uses one training seed and 100 epochs, and the evaluation data have informed development. The study motivates reporting query-level responses alongside aggregate metrics and identifying the intervention that supports each claim.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.