MLSTL-WSN: الكشف عن التسلل المعتمد على التعلم الآلي باستخدام SMOTETomek في الشبكات الاستشعارية اللاسلكية
MLSTL-WSN: machine learning-based intrusion detection using SMOTETomek in WSNs

المجلة: International Journal of Information Security، المجلد: 23، العدد: 3
DOI: https://doi.org/10.1007/s10207-024-00833-z
تاريخ النشر: 2024-03-19
المؤلف: Md. Alamin Talukder وآخرون
الموضوع الرئيسي: أمن الشبكات وكشف التسلل

نظرة عامة

في مجال الأنظمة السيبرانية الفيزيائية، تعتبر شبكات الاستشعار اللاسلكية (WSNs) بنى تحتية حيوية، تتكون من مستشعرات ثابتة ومتحركة تنظم نفسها بشكل مستقل لتشكيل روابط اتصال متعددة القفزات. هذه المستشعرات مسؤولة عن استشعار وجمع ومعالجة ونقل البيانات البيئية. ومع ذلك، فإن شبكات الاستشعار اللاسلكية عرضة لهجمات سريعة وضارة يمكن أن تؤثر على وظيفتها. تواجه طرق الكشف التقليدية عن التسلل تحديات كبيرة، بما في ذلك انخفاض معدلات الكشف، وارتفاع الأعباء الحاسوبية، والإنذارات الكاذبة، وذلك أساسًا بسبب قيود الموارد على عقد المستشعر، وتكرار البيانات، والارتباط العالي داخل الشبكة.

لمعالجة هذه التحديات، يقترح هذا البحث نهجًا جديدًا للكشف عن التسلل يدمج تقنيات التعلم الآلي مع خوارزمية تقنية زيادة العينة للأقليات الاصطناعية (SMOTE-TomekLink). تقوم هذه الطريقة بتوليف حالات الأقليات وإزالة روابط Tomek لإنشاء مجموعة بيانات متوازنة، مما يعزز دقة الكشف. تستخدم الدراسة توسيع الميزات من خلال التقييس لضمان ميزات إدخال متسقة وقابلة للتوسع، مما يسهل التدريب والكشف بدقة أكبر. من خلال تقييم النموذج على مجموعة بيانات شبكة الاستشعار اللاسلكية (WSN-DS) التي تحتوي على 374,661 سجلًا، تحقق النهج المقترح دقة مثيرة للإعجاب تبلغ 99.78% في التصنيف الثنائي و99.92% في سيناريوهات التصنيف المتعدد. بينما يظهر النموذج أداءً استثنائيًا، فإنه يواجه أيضًا قيودًا تتعلق بالتكاليف الحاسوبية والتأثيرات المحتملة من خيارات المعلمات الفائقة وجودة بيانات التدريب. تشمل اتجاهات البحث المستقبلية استكشاف طرق اختيار الميزات الهجينة، ونماذج التعلم العميق، والنهج الهرمية لتعزيز أداء الكشف عن التسلل في شبكات الاستشعار اللاسلكية.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على أهمية شبكات الاستشعار اللاسلكية (WSNs) كتكنولوجيا تحويلية لجمع البيانات ونقلها من عقد الاستشعار الموزعة عبر تطبيقات متنوعة، بما في ذلك مراقبة البيئة، والرعاية الصحية، والمدن الذكية. تؤكد الورقة على الأهمية الحرجة للأمان في شبكات الاستشعار اللاسلكية بسبب الطبيعة الحساسة للبيانات والثغرات الموجودة في هيكلها الموزع واللاسلكي. تحدد التهديدات الأمنية الرئيسية مثل الوصول غير المصرح به وتلاعب البيانات، مما يبرز ضرورة وجود آليات قوية للكشف عن التسلل.

يقترح المؤلفون نموذجًا جديدًا للكشف عن التسلل يدمج تقنيات التعلم الآلي (ML) مع تقنية زيادة العينة للأقليات الاصطناعية (SMOTE-Tomek) لمعالجة التحديات التي تطرحها مجموعات البيانات غير المتوازنة في شبكات الاستشعار اللاسلكية. يهدف هذا النموذج إلى تعزيز دقة وموثوقية الكشف عن التسلل من خلال توليف حالات الأقليات وتنقيح مجموعة البيانات. توضح الورقة استخدام خوارزميات التعلم الآلي المختلفة، بما في ذلك أشجار القرار، والغابات العشوائية، وطرق تعزيز التدرج، مما يمكّن النظام من تعلم الأنماط المعقدة والتكيف مع تقنيات الهجوم المتطورة. تشمل مساهمات هذا العمل نهجًا مبتكرًا للكشف عن التسلل، وتوسيع الميزات بشكل فعال، وتقييم شامل للنموذج، ومقارنات الأداء مع الطرق الحالية، مما يضع النموذج المقترح كتحسين قيم في أمان شبكات الاستشعار اللاسلكية.

طرق

في هذا القسم، يحدد المؤلفون منهجية شاملة لتعزيز الكشف عن التسلل في شبكات الاستشعار اللاسلكية (WSNs) من خلال دمج تقنيات التعلم الآلي (ML) مع خوارزمية SMOTE-Tomek. تبدأ الطريقة بجمع البيانات الخام من سيناريوهات WSN متنوعة لضمان القوة. تشمل خطوات المعالجة المسبقة توحيد ميزات الإدخال واستخدام ترميز التسمية للمتغيرات الفئوية، مما يعد البيانات لخوارزميات التعلم الآلي. لمعالجة مشكلة مجموعات البيانات غير المتوازنة، يتم تطبيق تقنية SMOTE-Tomek، مما يوازن بين حالات التسلل والحالات الطبيعية وبالتالي تحسين تعميم النموذج.

ثم يتم إخضاع مجموعة البيانات للتحقق المتقاطع باستخدام k-fold (10 طيات) لإنشاء مجموعات تدريب واختبار، مما يعزز التقييم الشامل لأداء النموذج. يتم تنفيذ العديد من خوارزميات التعلم الآلي، بما في ذلك الغابات العشوائية (RF)، وأشجار القرار (DT)، والشبكات العصبية متعددة الطبقات (MLP)، وأقرب الجيران (KNN)، وXGBoost (XGB)، وLight-GBM (LGB)، لالتقاط أنماط البيانات المختلفة. يستخدم تقييم النموذج مقاييس مثل الدقة، والدقة، والاسترجاع، ودرجة F1 لتحديد الفعالية في الكشف عن التسلل. يتم اختيار النموذج الأفضل أداءً بناءً على هذه التقييمات، مما يضمن ملاءمته للنشر في العالم الحقيقي. تتوج المنهجية بحل قوي لتعزيز أمان شبكات الاستشعار اللاسلكية، كما هو موضح في الأشكال والخوارزميات المرفقة.

نتائج

في التجارب التي أجريت على الكشف عن التسلل في شبكات الاستشعار اللاسلكية (WSNs)، تم إجراء تحليلات ثنائية ومتعددة التسميات، مما أظهر تحسينات كبيرة في كفاءة الكشف من خلال تطبيق تقنية توازن البيانات SMOTE-Tomek-Link (WiSTL). في النتائج الثنائية، حققت خوارزميات التعلم الآلي (ML) المختلفة، بما في ذلك أشجار القرار (DT)، والغابات العشوائية (RF)، والشبكات العصبية متعددة الطبقات (MLP)، وأقرب الجيران (KNN)، وLightGBM (LGB)، وXGBoost (XGB)، معدلات دقة عالية، حيث تفوقت RF باستمرار على الآخرين، محققة دقة تبلغ 99.78% وأدنى مقاييس خطأ (MAE، MSE، RMSE) عبر كلا السيناريوهين. أدى تضمين WiSTL إلى تحسين مقاييس الأداء بشكل عام، مما يبرز أهمية توازن البيانات في تعزيز فعالية النموذج.

في تحليل متعدد التسميات، لوحظت اتجاهات مماثلة، حيث أظهرت RF مرة أخرى أداءً متفوقًا، محققة دقة تبلغ 99.92% في سيناريو WiSTL. أشارت النتائج إلى أن توازن البيانات حسّن بشكل كبير من الدقة، والاسترجاع، ودرجات F1 لجميع النماذج. تم التحقق من أداء RF القوي من خلال معدلاته العالية من الإيجابيات الحقيقية والسلبيات الحقيقية، جنبًا إلى جنب مع الحد الأدنى من حالات الإيجابيات الكاذبة والسلبيات الكاذبة. وصلت المنطقة تحت منحنى ROC (AUC) لـ RF إلى 100%، مما يدل على قدرته الاستثنائية في الكشف بدقة عن التسللات مع الحفاظ على معدلات تصنيف خاطئة منخفضة. بشكل عام، تسلط النتائج الضوء على RF كأكثر الخوارزميات فعالية للكشف عن التسلل في شبكات الاستشعار اللاسلكية، مما يبرز ملاءمتها للتطبيقات الواقعية في تعزيز أمان الشبكة.

مناقشة

تسلط قسم المناقشة في الورقة البحثية الضوء على التقدم في أنظمة الكشف عن التسلل (IDS) لشبكات الاستشعار اللاسلكية (WSNs)، مع التركيز على فعالية نماذج التعلم الآلي (ML) المختلفة. بشكل ملحوظ، أظهر تان وآخرون فعالية خوارزمية الغابات العشوائية، محققين دقة تبلغ 92.57% بعد تطبيق تقنية زيادة العينة للأقليات الاصطناعية (SMOTE) لمعالجة عدم توازن مجموعة البيانات. وبالمثل، وجد ريزفي وآخرون أن الشبكات العصبية الاصطناعية (ANN) وأقرب الجيران (KNN) تميزت في الكشف عن هجمات حرمان الخدمة (DoS)، حيث حققت ANN دقة تبلغ 98.56%. قدم مينغ وآخرون طريقة تعتمد على LightGBM التي لم تحقق فقط معدلات كشف تزيد عن 99% ولكن أيضًا قللت بشكل كبير من وقت النمذجة من خلال تقليل أبعاد الميزات، مما يظهر إمكانيات تقنيات التعلم الآلي المتقدمة في تعزيز أداء IDS.

تناقش القسم أيضًا مجموعة من الأساليب المبتكرة، مثل الإطار متعدد الطبقات الذي قدمه الروحي وآخرون، والذي حقق معدلات دقة عالية عبر أنواع هجمات مختلفة، ونموذج التعلم العميق الذي قدمه تشاندري وآخرون، والذي وصل إلى معدل دقة يبلغ 97%. بالإضافة إلى ذلك، أظهرت الطريقة التجميعية التي قدمها دينر وآخرون أداءً استثنائيًا بدقة إجمالية بلغت 99.95%. تؤكد الورقة على أهمية معالجة عدم توازن الفئات من خلال تقنيات مثل SMOTE-TomekLink، التي حسنت بشكل فعال معدلات الكشف عبر فئات متعددة. بشكل عام، تشير النتائج إلى مسار واعد لنماذج الكشف عن التسلل المعتمدة على التعلم الآلي في شبكات الاستشعار اللاسلكية، مما يقترح أن الأساليب الهجينة قد تحقق دقة وموثوقية أكبر في الأبحاث المستقبلية.

Journal: International Journal of Information Security, Volume: 23, Issue: 3
DOI: https://doi.org/10.1007/s10207-024-00833-z
Publication Date: 2024-03-19
Author(s): Md. Alamin Talukder et al.
Primary Topic: Network Security and Intrusion Detection

Overview

In the realm of cyber-physical systems, wireless sensor networks (WSNs) serve as crucial infrastructures, comprising both stationary and mobile sensors that autonomously organize to form multi-hop communication links. These sensors are responsible for sensing, gathering, processing, and transmitting environmental data. However, WSNs are vulnerable to rapid and harmful attacks that can compromise their functionality. Traditional intrusion detection methods face significant challenges, including low detection rates, high computational overhead, and false alarms, primarily due to the resource constraints of sensor nodes, data redundancy, and high correlation within the network.

To address these challenges, this research proposes a novel intrusion detection approach that integrates machine learning techniques with the Synthetic Minority Oversampling Technique Tomek Link (SMOTE-TomekLink) algorithm. This method synthesizes minority instances and removes Tomek links to create a balanced dataset, thereby enhancing detection accuracy. The study employs feature scaling through standardization to ensure consistent and scalable input features, which facilitates more precise training and detection. Evaluating the model on the wireless sensor network dataset (WSN-DS) with 374,661 records, the proposed approach achieves an impressive accuracy of 99.78% in binary classification and 99.92% in multiclass classification scenarios. While the model demonstrates exceptional performance, it also faces limitations related to computational costs and potential influences from hyperparameter choices and training data quality. Future research directions include exploring hybrid feature selection methods, deep learning models, and hierarchical approaches to further enhance WSN intrusion detection performance.

Introduction

The introduction of this research paper highlights the significance of Wireless Sensor Networks (WSNs) as a transformative technology for data collection and transmission from distributed sensor nodes across various applications, including environmental monitoring, healthcare, and smart cities. The paper emphasizes the critical importance of security in WSNs due to the sensitive nature of the data and the vulnerabilities inherent in their distributed and wireless architecture. It identifies key security threats such as unauthorized access and data tampering, underscoring the necessity for robust intrusion detection mechanisms.

The authors propose a novel intrusion detection model that integrates Machine Learning (ML) techniques with the Synthetic Minority Oversampling Technique Tomek (SMOTE-Tomek) to address the challenges posed by imbalanced datasets in WSNs. This model aims to enhance the accuracy and reliability of intrusion detection by synthesizing minority instances and refining the dataset. The paper details the use of various ML algorithms, including Decision Trees, Random Forests, and gradient boosting methods, which enable the system to learn complex patterns and adapt to evolving attack techniques. The contributions of this work include an innovative approach to intrusion detection, effective feature scaling, comprehensive model evaluation, and performance comparisons with existing methods, positioning the proposed model as a valuable advancement in WSN security.

Methods

In this section, the authors outline a comprehensive methodology for enhancing intrusion detection in Wireless Sensor Networks (WSNs) by integrating machine learning (ML) techniques with the SMOTE-Tomek algorithm. The approach begins with the collection of raw data from diverse WSN scenarios to ensure robustness. Preprocessing steps include standardizing input features and employing label encoding for categorical variables, which prepares the data for ML algorithms. To tackle the issue of imbalanced datasets, the SMOTE-Tomek technique is applied, balancing normal and intrusion instances and thereby improving model generalization.

The dataset is then subjected to k-fold cross-validation (10 folds) to create training and testing sets, promoting a thorough assessment of model performance. Several ML algorithms, including Random Forest (RF), Decision Tree (DT), Multilayer Perceptron (MLP), K-Nearest Neighbors (KNN), XGBoost (XGB), and Light-GBM (LGB), are implemented to capture various data patterns. Model evaluation employs metrics such as accuracy, precision, recall, and F1-score to determine effectiveness in intrusion detection. The best-performing model is selected based on these evaluations, ensuring its suitability for real-world deployment. The methodology culminates in a robust solution for enhancing WSN security, as illustrated in the accompanying figures and algorithms.

Results

In the conducted experiments on intrusion detection in Wireless Sensor Networks (WSNs), both binary and multilabel analyses were performed, demonstrating significant improvements in detection efficiency through the application of the SMOTETomek-Link (WiSTL) data balancing technique. In the binary results, various machine learning (ML) algorithms, including Decision Trees (DT), Random Forest (RF), Multi-Layer Perceptron (MLP), K-Nearest Neighbors (KNN), LightGBM (LGB), and XGBoost (XGB), achieved high accuracy rates, with RF consistently outperforming others, achieving an accuracy of 99.78% and the lowest error metrics (MAE, MSE, RMSE) across both scenarios. The inclusion of WiSTL improved performance metrics across the board, underscoring the importance of data balancing in enhancing model efficacy.

In the multilabel analysis, similar trends were observed, with RF again demonstrating superior performance, achieving an accuracy of 99.92% in the WiSTL scenario. The results indicated that data balancing significantly improved the precision, recall, and F1 scores of all models. RF’s robust performance was further validated by its high true positive and true negative rates, alongside minimal false positive and false negative instances. The area under the ROC curve (AUC) for RF reached 100%, indicating its exceptional capability in accurately detecting intrusions while maintaining low misclassification rates. Overall, the findings highlight RF as the most effective algorithm for intrusion detection in WSNs, emphasizing its suitability for real-world applications in enhancing network security.

Discussion

The discussion section of the research paper highlights the advancements in intrusion detection systems (IDS) for wireless sensor networks (WSNs), emphasizing the effectiveness of various machine learning (ML) models. Notably, Tan et al. demonstrated the efficacy of the Random Forest algorithm, achieving an accuracy of 92.57% after applying the Synthetic Minority Oversampling Technique (SMOTE) to address dataset imbalance. Similarly, Rezvi et al. found that Artificial Neural Networks (ANN) and K-Nearest Neighbors (KNN) excelled in detecting Denial of Service (DoS) attacks, with ANN achieving an accuracy of 98.56%. Meng et al. introduced a LightGBM-based method that not only achieved over 99% detection rates but also significantly reduced modeling time through feature dimension reduction, showcasing the potential of advanced ML techniques in enhancing IDS performance.

The section further discusses various innovative approaches, such as the multi-tier framework by Alruhaily et al., which achieved high precision rates across different attack types, and the deep learning model by Chandre et al., which reached a 97% accuracy rate. Additionally, the ensemble approach by Dener et al. demonstrated exceptional performance with an overall accuracy of 99.95%. The paper underscores the importance of addressing class imbalance through techniques like SMOTE-TomekLink, which effectively improved detection rates across multiple classes. Overall, the findings indicate a promising trajectory for ML-based intrusion detection models in WSNs, suggesting that hybrid approaches may yield even greater accuracy and reliability in future research.