CPI: Contextual Prompt Injection against MCP Server Scanners
Abstract
Large language model (LLM)-based Model Context Protocol (MCP) Server scanners increasingly analyze not only executable code but also repository-level context, such as README files, tool descriptions, and code comments. Because this context can be freely manipulated by MCP maintainers, it creates an overlooked attack surface where a malicious maintainer may influence scanner judgments through security claims. Building on this insight, we introduce Contextual Prompt Injection (CPI), a black-box attack that uses a scanner’s own findings to construct vulnerability-specific remediation narratives, claiming that reported issues have been fixed while leaving the relevant executable implementation unchanged. To enable realistic evaluation, we curate a dataset containing 372 real-world MCP Servers, including 45 with verified vulnerabilities. We evaluate CPI against three LLM-based MCP security scanners and existing benchmarks. Evaluation with manual verification shows that CPI suppresses at least one genuine vulnerability finding in more than half of baseline-positive repositories for each evaluated scanner. Such findings demonstrate that LLM-based scanners can prioritize attacker-controlled remediation claims over contradictory code evidence, highlighting the need to treat repository context as untrusted input and independently verify claimed fixes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.