Visible Features, Hidden Flaws: Vulnerability Persistence in LLM-Generated Software
Abstract
Large language models LLMs increasingly drive automated multi-file software synthesis, yet frequently generate insecure backend implementations. Conventional code security evaluations treat generated artifacts as independent, post hoc evaluation samples, overlooking whether model generated failures share structured, reusable representations across generation tasks. In this work, we formalize and investigate vulnerability persistence: the hypothesis that generative coding models exhibit stable architectural biases that tightly couple user facing workflows with specific latent backend security flaws. We introduce the Feature-Security Table (), a model-specific empirical prior mapping standardized, observable user-interface interactions directly to ranked distributions of latent vulnerabilities without requiring access to victim source code, prompting history, or internal model weights. Evaluating six state-of-the-art code LLMs across multi-file applications in WebGenBench reveals strong vulnerability recurrence: recovers true latent vulnerabilities with up to 85.67% average attack success rate (ASR) and 84.97% average coverage (ACR) on held out application domains excluded entirely during construction, outperforming matched-budget baselines by up to 6.9. Notably, our recurrence framework demonstrates an unexpected universality gap cross domain vulnerability transfer consistently exceeds within domain recurrence by approximately 18 points proving that these security failures originate from model level coding habits rather than local prompt memorization or task specific artifacts. These findings establish that LLM-generated software exposes a predictable, observable attack surface, demonstrating that security evaluations must account for systematic, model level implementation biases rather than benchmarking generated programs independently. Our code is available at .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.