NeuroProbe: LLM-Augmented Reverse Engineering of Deep Neural Network Binaries
Abstract
Deep learning compilers translate high-level models into optimized native binaries yet these executables often leak structural patterns that expose proprietary architectures. We present NeuroProbe, a reverse-engineering pipeline that recovers operator semantics and architecture family from compiled neural network binaries under both clean and identifier-stripped conditions. Our solution ensembles a symbol-stripped XGBoost model trained on assembly features with a LoRA-tuned Qwen2.5-3B-Instruct model that outputs structured chain-of-thought rationale. We build and release a benchmark of 400 compiled binaries spanning three compiler backends (TVM, IREE, TensorRT) and three architecture families (CNN, Transformer, and Hybrid CNN-Transformer), covering 117 base models and 13,438 extracted kernel functions. NeuroProbe recovers operators with 99.7% F1 under clean compilation and 79.7% F1 under obfuscated setting, while architecture-family accuracy remains at 95% per binary in both settings. Beyond classification, NeuroProbe performs fine-grained hyperparameter estimation and topology reconstruction. We perform signal-ablation experiments to quantify the explicit contributions of symbol names versus assembly features, demonstrating that compiled transformer models remain substantially recoverable under heavy obfuscation despite prior assumptions of robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.