DOI: https://doi.org/10.1038/s41598-025-87454-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39870824
تاريخ النشر: 2025-01-27
المؤلف: Fatma S. Alrayes وآخرون
الموضوع الرئيسي: تقنيات الكشف عن الشذوذ وتطبيقاتها
نظرة عامة
تقدم ورقة البحث منهجية جديدة تُسمى التعلم الإحصائي مع الحفاظ على الخصوصية باستخدام خوارزمية تحسين لبيئة بيانات كبيرة عالية الأبعاد (PPSLOA-HDBDE). تتناول هذه الطريقة التحديات الكبيرة التي تطرحها زيادة عدد أجهزة إنترنت الأشياء (IoT)، التي تولد كميات هائلة من البيانات عالية الأبعاد بينما تثير مخاوف بشأن خصوصية البيانات وأمانها. تستخدم منهجية PPSLOA-HDBDE تقنيات تحسين متقدمة وتقنيات تجميع لضمان سرية البيانات دون المساس بالفعالية التحليلية. تشمل المكونات الرئيسية للمنهجية تطبيع القياس الخطي (LSN) لتهيئة البيانات، وعملية اختيار الميزات المعتمدة على مُحسن سرب القط الرملي (SCSO) للتخفيف من الأبعاد العالية، وتجميع من نماذج الشبكة العصبية التلافيفية الزمنية (TCN)، والمشفّر التلقائي متعدد الطبقات (MAE)، ونماذج تعزيز التدرج المتطرف (XGBoost) لاكتشاف التسلل.
تظهر عملية التحقق من أداء تقنية PPSLOA-HDBDE دقة ملحوظة تبلغ 99.49%، متفوقة على النماذج الحالية. ومع ذلك، تواجه المنهجية قيودًا، مثل إمكانية تباين البيانات عبر أجهزة إنترنت الأشياء المتنوعة، مما قد يؤثر على دقة نتائج التعلم الفيدرالي. بالإضافة إلى ذلك، قد تتحدى المطالب الحاسوبية للتحليل في الوقت الحقيقي الأجهزة الطرفية ذات الموارد المحدودة، وتستمر المخاوف المتعلقة بخصوصية مشاركة البيانات. تشمل اتجاهات البحث المستقبلية استكشاف نماذج محسّنة للحفاظ على الخصوصية، وآليات تكيفية لتدريب النماذج بناءً على قدرات الأجهزة المتنوعة، وتوسيع تطبيق النموذج ليشمل مجالات أوسع تتجاوز الإلكترونيات الاستهلاكية. كما يُقترح دمج ملاحظات المستخدمين في عملية التوصية لتحسين التخصيص ورضا المستخدم.
طرق
تقدم البحث منهجية PPSLOA-HDBDE، المصممة لتعزيز سرية البيانات مع ضمان الفعالية التحليلية من خلال تقنيات تحسين متقدمة وتقنيات تجميع. تتكون المنهجية من أربع عمليات رئيسية: تهيئة البيانات، اختيار الميزات باستخدام SCSO، نماذج التصنيف التجميعية، وتحسين المعلمات عبر IMPA. يتم التحقق من أداء هذا النموذج باستخدام مجموعة بيانات BoT-IoT، التي تشمل 2056 عينة عبر خمس فئات، مما يحاكي بيئة شبكة واقعية مع أنواع مختلفة من الهجمات، بما في ذلك DDoS وDoS.
تشير النتائج التجريبية إلى أن منهجية PPSLOA-HDBDE تحقق دقة عالية في تصنيف جميع الخمس تسميات، مع ملاحظة مقاييس أداء ملحوظة عند عدد مختلف من العصور. على وجه التحديد، عند 500 عصر، يُبلغ النموذج عن دقة تبلغ 98.89%، ودقة 96.43%، واسترجاع 96.64%، ودرجة F1 تبلغ 96.53%، وMCC تبلغ 95.82%. تتحسن هذه المقاييس قليلاً عند 1000 عصر، مع دقة تبلغ 99.11%. تُظهر اتجاهات دقة التدريب والتحقق الحد الأدنى من الإفراط في التكيف، بينما تؤكد تحليلات دقة الاسترجاع ومنحنى ROC قدرات التصنيف القوية للنموذج. بالمقارنة، يتفوق PPSLOA-HDBDE على الطرق الحالية، محققًا دقة تبلغ 99.49%، متجاوزًا العديد من النماذج المعاصرة.
مناقشة
تستعرض قسم المناقشة في ورقة البحث مختلف التقدمات في أنظمة اكتشاف ومنع التسلل (IDPS) عبر سياقات متعددة، مع تسليط الضوء على دمج الذكاء الاصطناعي (AI)، وتعلم الآلة (ML)، وتقنيات البلوكشين. تشمل المساهمات الملحوظة استخدام الخوارزميات الجينية (GA) لتحسين التوجيه في الشبكات اللاسلكية المتنقلة (MWSNs) لتعزيز كفاءة الطاقة والموثوقية. بالإضافة إلى ذلك، تم تطوير طرق قائمة على الذكاء الاصطناعي لأنظمة اكتشاف التسلل (IDS) لتصنيف الهجمات واستخدام نظرية الألعاب للتوجيه الآمن، بينما تهدف استراتيجيات جديدة مثل أساليب الفخاخ والنهج المعتمدة على البلوكشين إلى تقليل زمن النقل.
تشدد الورقة على أهمية تقنيات اختيار الميزات وتقليل الأبعاد، مثل نموذج تحسين بحث القط الرملي (SCSO)، الذي يعالج بفعالية تحديات البيانات عالية الأبعاد من خلال تحقيق توازن بين الاستكشاف والاستغلال. علاوة على ذلك، يتم اقتراح نماذج التصنيف التجميعية التي تجمع بين الشبكات التلافيفية الزمنية (TCN)، ومتوسط الخطأ المطلق (MAE)، وXGBoost للاستفادة من نقاط قوتها التكميلية، مما يعزز دقة الاكتشاف ومرونة النموذج ضد أنواع الهجمات المتنوعة. على الرغم من هذه التقدمات، يحدد المؤلفون فجوة بحثية كبيرة في تطوير حلول قابلة للتكيف وقابلة للتوسع للبيئات في الوقت الحقيقي، مما يبرز الحاجة إلى نهج شامل لمعالجة تعقيدات التهديدات السيبرانية الحديثة بشكل فعال.
DOI: https://doi.org/10.1038/s41598-025-87454-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39870824
Publication Date: 2025-01-27
Author(s): Fatma S. Alrayes et al.
Primary Topic: Anomaly Detection Techniques and Applications
Overview
The research paper introduces a novel methodology termed Privacy-Preserving Statistical Learning with an Optimization Algorithm for a High-Dimensional Big Data Environment (PPSLOA-HDBDE). This approach addresses the significant challenges posed by the proliferation of Internet of Things (IoT) devices, which generate vast amounts of high-dimensional data while raising concerns about data privacy and security. The PPSLOA-HDBDE methodology employs advanced optimization and ensemble techniques to ensure data confidentiality without compromising analytical efficacy. Key components of the methodology include linear scaling normalization (LSN) for data preprocessing, a sand cat swarm optimizer (SCSO)-based feature selection process to mitigate high dimensionality, and an ensemble of temporal convolutional network (TCN), multi-layer auto-encoder (MAE), and extreme gradient boosting (XGBoost) models for intrusion detection.
The performance validation of the PPSLOA-HDBDE technique demonstrates a remarkable accuracy of 99.49%, outperforming existing models. However, the methodology faces limitations, such as potential data heterogeneity across diverse IoT devices, which may impact the accuracy of federated learning results. Additionally, the computational demands of real-time analysis could challenge resource-constrained edge devices, and privacy concerns regarding data sharing persist. Future research directions include exploring enhanced privacy-preserving models, adaptive mechanisms for model training based on varying device capabilities, and expanding the applicability of the model to broader domains beyond consumer electronics. Incorporating user feedback into the recommendation process is also suggested to improve personalization and user satisfaction.
Methods
The research introduces the PPSLOA-HDBDE methodology, designed to enhance data confidentiality while ensuring analytical effectiveness through advanced optimization and ensemble techniques. The methodology comprises four key processes: data preprocessing, feature selection using SCSO, ensemble classification models, and parameter optimization via IMPA. The performance of this model is validated using the BoT-IoT dataset, which includes 2056 samples across five classes, simulating a realistic network environment with various attack types, including DDoS and DoS.
Experimental results indicate that the PPSLOA-HDBDE methodology achieves high accuracy in classifying all five labels, with notable performance metrics observed at different epoch counts. Specifically, at 500 epochs, the model reports an accuracy of 98.89%, precision of 96.43%, recall of 96.64%, F1 score of 96.53%, and MCC of 95.82%. These metrics improve slightly at 1000 epochs, with an accuracy of 99.11%. The training and validation accuracy trends demonstrate minimal overfitting, while the precision-recall and ROC curve analyses further confirm the model’s robust classification capabilities. Comparatively, the PPSLOA-HDBDE outperforms existing methods, achieving an accuracy of 99.49%, surpassing several contemporary models.
Discussion
The discussion section of the research paper reviews various advancements in intrusion detection and prevention systems (IDPS) across multiple contexts, highlighting the integration of artificial intelligence (AI), machine learning (ML), and blockchain technologies. Notable contributions include the use of genetic algorithms (GA) for optimizing routing in mobile wireless sensor networks (MWSNs) to enhance energy efficiency and reliability. Additionally, AI-based methods for intrusion detection systems (IDS) have been developed to classify attacks and utilize game theory for secure routing, while novel strategies such as honeypot and blockchain-based approaches aim to reduce transmission latency.
The paper emphasizes the importance of feature selection and dimensionality reduction techniques, such as the Sand Cat Search Optimization (SCSO) model, which effectively addresses high-dimensional data challenges by balancing exploration and exploitation. Furthermore, ensemble classification models combining Temporal Convolutional Networks (TCN), Mean Absolute Error (MAE), and XGBoost are proposed to leverage their complementary strengths, enhancing detection accuracy and resilience against diverse attack types. Despite these advancements, the authors identify a significant research gap in developing adaptable and scalable solutions for real-time environments, underscoring the need for holistic approaches to address the complexities of modern cyber threats effectively.
