Donguk Min, Seungsoo Nam, Daeseon Choi · Journal of Information Security and Applications 2026 · 2026
DOI: 10.1016/j.jisa.2026.104652
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Machine-learning-based network intrusion detection systems (ML-NIDS) are now widely deployed as security services, yet they remain vulnerable to model extraction. Prior work typically relies on the soft-label confidence scores of the victim, whereas ML-NIDS in production usually return only the Top-1 prediction. Network traffic is also tabular and severely class-imbalanced, so techniques developed for image models do not transfer directly. We show that diverse ML-NIDS can be cloned with high fidelity from hard-label queries combined with a small proxy set, without any internal information or confidence scores. To this end, we propose a framework that jointly trains a clone model and a sample generator, so that the generator produces informative queries and adapts to the characteristics of tabular network data. The framework further stabilizes training and preserves high extraction quality on rare attack classes, which are easily degraded under conventional optimization. We evaluate the attack under realistic black-box conditions on NSL-KDD, CIC-IDS2017, and CSE-CIC-IDS2018, including a cross-dataset distribution shift. Across five architecturally distinct ML-NIDS on CIC-IDS2017, it attains 82–99% Mean Per-Class Fidelity under a limited query budget, showing that restricting API outputs to Top-1 predictions alone is insufficient to protect deployed ML-NIDS against model extraction.
No comments yet — start the discussion below.