نظام كشف التسلل القائم على التعلم الآلي القوي باستخدام تقنيات إحصائية بسيطة في اختيار الميزات
Robust machine learning based Intrusion detection system using simple statistical techniques in feature selection

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-88286-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39893240
تاريخ النشر: 2025-02-01
المؤلف: Sunil Kaushik وآخرون
الموضوع الرئيسي: أمن الشبكات وكشف التسلل

نظرة عامة

إن الانتشار السريع لأجهزة إنترنت الأشياء (IoT)، الضرورية لتقدم الصناعة 4.0، قد أثار مخاوف أمنية كبيرة. تعمل هذه الأجهزة غالبًا تحت قيود الطاقة المحدودة وقدرات المعالجة، مما يجعلها عرضة للهجمات الإلكترونية ويعقد نشر أنظمة الكشف عن التسلل الفعالة (IDS). لمواجهة هذه التحديات، يقدم البحث خوارزمية مبتكرة لاختيار الميزات تعتمد على طرق إحصائية أساسية، إلى جانب IDS خفيف الوزن. لا تعزز هذه الطريقة الأداء فحسب، بل تقلل أيضًا من وقت التدريب بنسبة 27-63% عبر مصنفات مختلفة. من خلال التركيز على الميزات الأكثر تمييزًا، تقلل المنهجية المقترحة من الحمل الحاسوبي بينما تحسن بشكل كبير من دقة الكشف.

أظهر IDS أداءً استثنائيًا، محققًا أكثر من 99.9% في الدقة، والدقة، والاسترجاع، وF1-Score على مجموعة بيانات IoTID20، مع نتائج متسقة على مجموعة بيانات NSLKDD. تؤكد هذه النتائج فعالية خوارزمية اختيار الميزات المقترحة وIDS في معالجة الثغرات الأمنية المرتبطة بأجهزة إنترنت الأشياء. من المحتمل أن تستكشف الأعمال المستقبلية المزيد من التحسينات وتطبيق هذه المنهجية على مجموعة أوسع من بيئات إنترنت الأشياء.

مقدمة

تقدم مقدمة ورقة البحث معلومات أساسية ضرورية ذات صلة بالدراسة. توضح المفاهيم والنظريات الرئيسية التي تدعم البحث، مما يضع أساسًا للتحليل اللاحق. يبرز المؤلفون أهمية الموضوع ضمن الخطاب الأكاديمي الأوسع، مؤكدين على صلته بالتحديات الحالية أو الفجوات في المعرفة.

علاوة على ذلك، تهيئ المقدمة المسرح لأهداف البحث والأسئلة، مشيرة إلى الأهداف المحددة للدراسة. من خلال وضع البحث في سياق الأدبيات الموجودة، يظهر المؤلفون ضرورة تحقيقهم وإسهاماته المحتملة في المجال. بشكل عام، تعمل هذه القسم على جذب القارئ وتوضيح أهمية البحث الذي تم القيام به.

طرق

تصف منهجية البحث الموضحة في هذا القسم الإطار التجريبي المستخدم للتحقيق في الكشف عن التسلل الشبكي باستخدام تقنيات التعلم الآلي. استخدمت الدراسة PySpark على منصة Google Colab، مستفيدة من Scikit-learn للتعلم الآلي واختيار الميزات. شملت الإعدادات الحاسوبية معالج Intel Core i5-1345U من الجيل الثالث عشر مع 8 أنوية، و16 جيجابايت من ذاكرة الوصول العشوائي، و512 جيجابايت SSD، مما يسهل تدريب النماذج واختبارها بشكل قوي. تم إجراء التجارب على مدى عشر تكرارات، مع إجراء تحليل إحصائي باستخدام اختبار t. تم التلاعب بمجموعة البيانات الرئيسية، IoTID20، لمحاكاة هجمات يوم الصفر عن طريق تقسيمها إلى مجموعات تدريب حميدة ومجموعات اختبار خبيثة. بالإضافة إلى ذلك، تم استخدام مجموعة بيانات NSLKDD لتقييم قابلية توسيع تقنية اختيار الميزات المقترحة عبر أحجام مجموعات بيانات مختلفة.

شملت المنهجية تدريب مصنفات معقدة مثل Random Forest (RF) وAdaBoost (AD)، إلى جانب خوارزميات أبسط مثل الانحدار اللوجستي (LR)، وآلة الدعم الناقل (SVM)، وNaive Bayes (NB)، وأشجار القرار (DT). تم تحويل مجموعات البيانات إلى ثنائية للتصنيف، وتم تقييم مقاييس الأداء مثل الدقة (P_A)، والدقة (P_R)، ومعدل الإيجابيات الكاذبة (P_F)، والأداء التنبؤي (P_p). تم تحديد المصنفات ذات الأداء الضعيف (LR، SVM، NB) وخضعت لاختيار الميزات باستخدام طرق مثل مربع كاي، وارتباط بيرسون، وANOVA، مما أسفر عن مجموعة بيانات مصقولة تسمى مجموعة بيانات Chico، التي تضمنت 17 ميزة لـ IoTID20 و16 لـ NSLKDD. ثم تم تطبيق المصنفات عالية الأداء (DT، RF، AD) على هذه المجموعة، مع مقارنة أدائها ضد الطرق المعاصرة. تم إجراء تحسين النموذج باستخدام GridSearchCV لتعزيز الأداء من خلال ضبط المعلمات، كما هو موضح في الخوارزميات والأشكال المرفقة.

نتائج

تظهر نتائج الدراسة فعالية خوارزميات التعلم الآلي (ML) المختلفة، بما في ذلك المصنفات الأساسية (الانحدار اللوجستي، Naive Bayes، آلة الدعم الناقل، شجرة القرار) وطرق التجميع (Random Forest، AdaBoost)، في تصنيف الهجمات على مجموعات بيانات IoTID20 وNSLKDD. كشفت التجارب الأولية أن المصنفات الأساسية حققت دقة تتراوح بين 60% و70%، مما يشير إلى كفاءة محدودة في تصنيف الهجمات. تم استخدام تقنيات اختيار الميزات، وخاصة مربع كاي (CHI)، وارتباط بيرسون، وANOVA، لتعزيز أداء المصنف. أظهر التحليل أن المصنفات أدت بشكل مثالي مع 10 إلى 15 ميزة، حيث حقق Naive Bayes أعلى دقة تبلغ حوالي 75% باستخدام أفضل 15 ميزة تم اختيارها بواسطة إما طرق CHI أو بيرسون.

أدى المزيد من التحسين إلى إنشاء مجموعة بيانات جديدة، تسمى مجموعة بيانات Corchi، التي تضمنت 26 ميزة فريدة تم اختيارها من مجموعة بيانات IoTID20. حسنت هذه المجموعة بشكل كبير من أداء المصنف، حيث حققت شجرة القرار دقة تبلغ 99.8%، بينما وصلت Random Forest وAdaBoost إلى حوالي 98%. أكدت التحقق مع مجموعة بيانات NSLKDD هذه النتائج، حيث تفوقت شجرة القرار مرة أخرى على المصنفات الأخرى بدقة 99.9%. أكد التحليل الإحصائي أهمية النتائج، مع قيم p أقل من 0.01، مما يبرز قوة تقنيات اختيار الميزات المقترحة في تعزيز دقة أنظمة الكشف عن التسلل.

مناقشة

تؤكد قسم المناقشة في الورقة على الدور الحاسم لأنظمة الكشف عن التسلل (IDSs) في التخفيف من التهديد المتزايد للهجمات الإلكترونية على المنظمات الحديثة. تصنف IDSs إلى أنظمة قائمة على التوقيع وأنظمة قائمة على الشذوذ، مع تسليط الضوء على قيود الأولى، التي تعتمد على توقيعات الهجمات المحددة مسبقًا وتتطلب تحديثات مستمرة. في المقابل، يمكن لأنظمة IDS القائمة على الشذوذ، المعززة بتقنيات التعلم الآلي (ML)، تحديد التهديدات غير المرئية سابقًا بشكل تكيفي. تشير مراجعة الأدبيات إلى اتجاه متزايد نحو دمج التعلم العميق وطرق التجميع، التي أظهرت أداءً متفوقًا من حيث الدقة وتقليل معدلات التصنيف الخاطئ مقارنة بخوارزميات التعلم الآلي السطحية التقليدية.

يقدم البحث نهجًا جديدًا يستخدم طرق اختيار الميزات الإحصائية الأساسية مع مصنفات التعلم الآلي البسيطة، مثل أشجار القرار وGaussian Naïve Bayes، لتطوير IDS خفيف الوزن مصمم لإنترنت الأشياء (IoT). يظهر النظام المقترح دقة تنافسية (99.89%) ودقة (99.86%) مع الحفاظ على تعقيد حسابي أقل مقارنة بالنماذج الحالية التي تستخدم تقنيات أكثر تعقيدًا. يدعو المؤلفون إلى إجراء أبحاث مستقبلية لاستكشاف التعلم الانتقالي والشبكات العصبية العميقة، بهدف تحقيق توازن بين دقة الكشف العالية واستخدام موارد الحوسبة بكفاءة، مما يعزز قدرات الكشف عن التسلل في الوقت الحقيقي.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-88286-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39893240
Publication Date: 2025-02-01
Author(s): Sunil Kaushik et al.
Primary Topic: Network Security and Intrusion Detection

Overview

The rapid proliferation of Internet of Things (IoT) devices, crucial for the advancement of Industry 4.0, has raised significant security concerns. These devices often operate under constraints of limited energy and processing capabilities, rendering them vulnerable to cyberattacks and complicating the deployment of effective intrusion detection systems (IDS). To tackle these challenges, the study introduces an innovative feature selection algorithm grounded in basic statistical methods, coupled with a lightweight IDS. This approach not only enhances performance but also reduces training time by 27-63% across various classifiers. By focusing on the most discriminative features, the proposed methodology minimizes computational overhead while significantly improving detection accuracy.

The IDS demonstrated exceptional performance, achieving over 99.9% in accuracy, precision, recall, and F1-Score on the IoTID20 dataset, with consistent results on the NSLKDD dataset. These findings underscore the effectiveness of the proposed feature selection algorithm and IDS in addressing the security vulnerabilities associated with IoT devices. Future work will likely explore further optimizations and the application of this methodology to a broader range of IoT environments.

Introduction

The introduction of the research paper provides essential background information relevant to the study. It outlines the key concepts and theories that underpin the research, establishing a foundation for the subsequent analysis. The authors highlight the significance of the topic within the broader academic discourse, emphasizing its relevance to current challenges or gaps in knowledge.

Furthermore, the introduction sets the stage for the research objectives and questions, indicating the specific aims of the study. By contextualizing the research within existing literature, the authors demonstrate the necessity of their investigation and its potential contributions to the field. Overall, this section serves to engage the reader and clarify the importance of the research undertaken.

Methods

The research methodology outlined in this section describes the experimental framework employed to investigate network intrusion detection using machine learning techniques. The study utilized PySpark on the Google Colab platform, leveraging Scikit-learn for machine learning and feature selection. The computational setup included a 13th generation Intel Core i5-1345U processor with 8 cores, 16 GB of RAM, and a 512 GB SSD, facilitating robust model training and testing. Experiments were conducted over ten iterations, with statistical analysis performed using the t-test. The primary dataset, IoTID20, was manipulated to simulate zero-day attacks by segmenting it into benign training and malicious testing sets. Additionally, the NSLKDD dataset was employed to assess the scalability of the proposed feature selection technique across varying dataset sizes.

The methodology involved training complex classifiers such as Random Forest (RF) and AdaBoost (AD), alongside simpler algorithms like Logistic Regression (LR), Support Vector Machine (SVM), Naive Bayes (NB), and Decision Trees (DT). The datasets were binarized for classification, and performance metrics such as accuracy (P_A), precision (P_R), false positive rate (P_F), and predictive performance (P_p) were evaluated. Poorly performing classifiers (LR, SVM, NB) were identified and subjected to feature selection using methods like chi-square, Pearson correlation, and ANOVA, resulting in a refined dataset termed the Chico dataset, which comprised 17 features for IoTID20 and 16 for NSLKDD. High-performing classifiers (DT, RF, AD) were then applied to this dataset, with their performance compared against contemporary methods. The model optimization was conducted using GridSearchCV to enhance performance through hyperparameter tuning, as detailed in the accompanying algorithms and figures.

Results

The results of the study demonstrate the effectiveness of various machine learning (ML) algorithms, including basic classifiers (Logistic Regression, Naive Bayes, Support Vector Machine, Decision Tree) and ensemble methods (Random Forest, AdaBoost), in classifying attacks on IoTID20 and NSLKDD datasets. Initial experiments revealed that the basic classifiers achieved accuracies between 60% and 70%, indicating limited efficiency in attack classification. Feature selection techniques, specifically Chi-square (CHI), Pearson correlation, and ANOVA, were employed to enhance classifier performance. The analysis showed that classifiers performed optimally with 10 to 15 features, with Naive Bayes achieving the highest accuracy of approximately 75% using the top 15 features selected by either CHI or Pearson methods.

Further refinement led to the creation of a new dataset, termed the Corchi dataset, which comprised 26 unique features selected from the IoTID20 dataset. This dataset significantly improved classifier performance, with Decision Tree achieving an accuracy of 99.8%, while Random Forest and AdaBoost reached approximately 98%. Validation with the NSLKDD dataset corroborated these findings, with Decision Tree again outperforming other classifiers at 99.9% accuracy. Statistical analysis confirmed the significance of the results, with p-values below 0.01, underscoring the robustness of the proposed feature selection techniques in enhancing the accuracy of intrusion detection systems.

Discussion

The discussion section of the paper emphasizes the critical role of Intrusion Detection Systems (IDSs) in mitigating the increasing threat of cyberattacks on modern organizations. It categorizes IDSs into signature-based and anomaly-based systems, highlighting the limitations of the former, which relies on predefined attack signatures and requires constant updates. In contrast, anomaly-based IDSs, enhanced by machine learning (ML) techniques, can adaptively identify previously unseen threats. The literature review indicates a growing trend towards integrating deep learning and ensemble methods, which have shown superior performance in terms of accuracy and reduced misclassification rates compared to traditional shallow ML algorithms.

The study presents a novel approach that utilizes basic statistical feature selection methods combined with simple machine learning classifiers, such as Decision Trees and Gaussian Naïve Bayes, to develop a lightweight IDS tailored for the Internet of Things (IoT). The proposed system demonstrates competitive accuracy (99.89%) and precision (99.86%) while maintaining lower computational complexity compared to existing models that employ more sophisticated techniques. The authors advocate for future research to explore transfer learning and deep neural networks, aiming to balance high detection accuracy with efficient computational resource utilization, thereby enhancing real-time intrusion detection capabilities.

شارك: