acceptodds
Under review as a conference paper at ICLR 2027

VANISH: Fusing Source Code and Virtual Assembly for Robust Cross-Language Code Clone Detection

Abstract

Cross-language code clone detection supports code search, migration, and plagiarism analysis, yet remains vulnerable to semantics-preserving obfuscation because existing models rely heavily on lexical evidence. We propose VANISH, a dual-view framework that combines multilingual source code with identifier-reduced virtual assembly. VANISH resolves supervision, tokenization, and granularity mismatches, trains both views independently with cross-language contrastive learning, and fuses their similarity scores. We also construct XL-Asm, a leakage-controlled benchmark of 91,863 programs from 29,324 functionality groups across 12 languages. VANISH achieves state-of-the-art performance with 0.674 X-MRR, outperforming CodeSage and UniXcoder by 12.3% and 68.5%, respectively. Under strong obfuscation, it recovers detection AUC from 0.85–0.88 to 0.94–0.95, gaining up to 8.9 percentage points (*p* < 0.001) without reducing clean-code performance. Results across languages, source encoders, real obfuscators, and adaptive attacks establish the *headroom principle*, whereby virtual assembly contributes most when source representations are degraded.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.