CuE: A Cue Engineering Framework for Diagnosing and Mitigating Shortcut Effects on Zero-shot MLLM Recognition
Abstract
Generative multimodal large language models (MLLMs) offer open-vocabulary recognition, but their decisions can depend on spurious visual or linguistic cues. We introduce CuE, a training-free framework that organizes candidate cue mechanisms, constructs target-conditioned cue hypotheses with an offline helper LLM, and performs fixed-label diagnosis before cue-matched remediation on frozen MLLMs based on a taxonomy of cue families introduced in this study. On a stratified 12-target COCO panel, we evaluate four target MLLMs from three model families. Visual diagnosis reveals recurring sensitivities: target recall is %–% lower in competitor-present groups, and partial-evidence groups expose %–% recall gaps. Language-side diagnosis further shows model-dependent sensitivities: query rewrite flips %–% of decisions, while language-prior probes and text interference increase miss rates by %–% and %–%, respectively. Cue-matched remediation improves balanced accuracy consistently across the evaluated visual settings, whereas generic debiasing baselines can over-constrain recognition and reduce recall. These findings suggest that robust zero-shot MLLM recognition requires diagnosing the relevant shortcut cue families before applying remediation, rather than relying on one-size-fits-all prompting fixes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.