Can MLLMs Assist Forensic Sketch Identification via Semantic Reasoning?
Abstract
When an eyewitness describes a suspect, the description is conveyed through language rather than pixels. Forensic face recognition, however, has largely been studied as a purely visual task, and the semantic content of verbal, written, and spoken facial descriptions has received little attention. We ask whether Multimodal Large Language Models (MLLMs) can bridge language and vision by jointly reasoning over facial sketches, textual descriptions, and spoken biometric narratives for forensic identity recognition. To study this, we introduce MS-Forensics, the first multimodal benchmark for sketch-based forensic identity recognition, which pairs facial sketches with textual and audio-based biometric descriptions built from five public face and sketch datasets. Through hypothesis-driven evaluation, we show that recognition degrades as the candidate set grows, and that the saturation of description-based recognition is explained by a collision model over attribute combinations. Fusing sketches with textual or audio descriptions outperforms unimodal inputs in most settings, but MLLMs struggle under fine-grained identity similarity and rely heavily on demographic cues; once demographics are removed, the eyes and forehead emerge as the most discriminative attributes. MS-Forensics provides a benchmark for evaluating MLLMs in AI-assisted forensic analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.