VANISH: Fusing Source Code and Virtual Assembly for Robust Cross-Language Code Clone Detection
Abstract
Cross-language code clone detection supports code search, migration, and plagiarism analysis, yet remains vulnerable to semantics-preserving obfuscation because existing models rely heavily on lexical evidence. We propose VANISH, a dual-view framework that combines multilingual source code with identifier-reduced virtual assembly. VANISH resolves supervision, tokenization, and granularity mismatches, trains both views independently with cross-language contrastive learning, and fuses their similarity scores. We also construct XL-Asm, a leakage-controlled benchmark of 91,863 programs from 29,324 functionality groups across 12 languages. VANISH achieves state-of-the-art performance with 0.674 X-MRR, outperforming CodeSage and UniXcoder by 12.3% and 68.5%, respectively. Under strong obfuscation, it recovers detection AUC from 0.85–0.88 to 0.94–0.95, gaining up to 8.9 percentage points (*p* < 0.001) without reducing clean-code performance. Results across languages, source encoders, real obfuscators, and adaptive attacks establish the *headroom principle*, whereby virtual assembly contributes most when source representations are degraded.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.