اختيار الميزات الأمثل لتوقع أمراض القلب باستخدام خوارزمية مستعمرة النحل الاصطناعية المعدلة (M-ABC) وجيران K الأقرب (KNN)
Optimal feature selection for heart disease prediction using modified Artificial Bee colony (M-ABC) and K-nearest neighbors (KNN)

شارك:
المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-78021-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39482391
تاريخ النشر: 2024-10-31
المؤلف: Muhammad Amir Khan وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تقدم الورقة البحثية نهجًا جديدًا للتشخيص المبكر لأمراض القلب باستخدام نموذج تصنيف يدمج خوارزمية مستعمرة النحل الاصطناعية المعدلة (M-ABC) مع أقرب الجيران (KNN) لاختيار الميزات المثلى. تؤكد الدراسة على أهمية تحديد السمات المهمة من مجموعة البيانات لتعزيز دقة التصنيف وتقليل وقت التدريب. من خلال استخدام خوارزمية النحل المعدلة، يحتفظ الإطار المقترح فقط بأكثر الميزات إفادة خلال مرحلة تدريب التصنيف، مما يسهل التنبؤات الأكثر دقة واتخاذ قرارات مستنيرة في بيئات الرعاية الصحية.

في الختام، يبرز المؤلفون فعالية نموذجهم المقترح، الذي يتفوق على الخوارزميات الأساسية، بما في ذلك الخوارزمية الجينية مع آلة الدعم الناقل (SVM) وKNN، في مقاييس الأداء المختلفة مثل الدقة والحساسية والاسترجاع وF-score والدقة والخصوصية. على وجه التحديد، يظهر النموذج تحسينات بنسبة 2.6% في الدقة ومقاييس أخرى، مما يدل على إمكانيته للتطبيق العملي في التنبؤ بأمراض القلب مع معالجة مخاوف الخصوصية. ومع ذلك، يعترف المؤلفون بالقيود، مثل التحديات في التعامل مع مواقع عملاء إنترنت الأشياء الطبية (IoMT) المتعددة وحساسية النموذج لتعديلات معدل التعلم. يقترحون أن الأبحاث المستقبلية يمكن أن تستكشف تقنيات التعلم الآلي المتقدمة، بما في ذلك نماذج التعلم العميق، لتعزيز قدرات التنبؤ بمختلف الأمراض الخطيرة.

الطرق

تقدم الورقة البحثية إطارًا للتعلم الفيدرالي مصممًا للتنبؤ بأمراض القلب ضمن نظام رعاية صحية واعٍ للخصوصية يستخدم إنترنت الأشياء الطبية (IoMT). يعالج هذا الإطار التحديات المتعلقة بالخصوصية التي تفرضها اللوائح مثل GDPR وقانون الأمن السيبراني في الصين وCCPA، والتي تعقد الأساليب التقليدية للتعلم الآلي التي تتطلب معالجة بيانات المستخدم الحساسة. تستخدم الدراسة مجموعة بيانات أمراض القلب في UCI كليفلاند، التي تتكون من 303 سجلات و14 سمة، لتدريب وتقييم النموذج المقترح. تم التعامل مع القيم المفقودة للميزات المحددة باستخدام تعويض المتوسط، بينما تم ترميز الميزات الفئوية من خلال الترميز الأحادي.

للتقييم التجريبي، أجرى المؤلفون محاكاة على مدى أربعة آلاف جولة باستخدام بيئة بايثون مع PyTorch. تم تنفيذ المحاكاة على معالج Intel® Core™ i7-8550 بسرعة 4 جيجاهرتز، مما يضمن اتساق الإعداد التجريبي. شملت خطوات معالجة البيانات التنظيف والتوحيد والترميز واختيار الميزات باستخدام MABC لتقليل الأبعاد. للتخفيف من عدم توازن الفئات، تم تطبيق تقنية الزيادة الاصطناعية للأقليات (SMOTE). تدعم هذه المنهجية الشاملة الادعاءات المتعلقة بمتانة وفعالية الإطار المقترح في التنبؤ بأمراض القلب، كما يتضح من دقة الإبلاغ والدقة والاسترجاع ومقاييس F1-score.

النتائج

يظهر الإطار المقترح للتنبؤ بأمراض القلب أداءً متفوقًا مقارنةً بالتقنيات الحالية، بما في ذلك الخوارزميات الجينية المدمجة مع آلات الدعم الناقل (SVM) وأقرب الجيران (kNN). على وجه التحديد، يحقق الإطار حساسية بنسبة 85.6%، مما يدل على قدرته القوية على تحديد الحالات الإيجابية بدقة لأمراض القلب. تتجاوز هذه الحساسية تلك الخاصة بالطرق الأخرى التي تم تقييمها، كما هو موضح في الجدول 3 والشكل 4.

بالإضافة إلى ذلك، يظهر الإطار معدل استرجاع بنسبة 84.5%، مما يعكس فعاليته في التقاط جميع حالات النتائج المواتية. يتفوق هذا المعدل أيضًا على تلك الخاصة بالخوارزميات الأخرى التي تم تقييمها. تسلط النتائج المقارنة الضوء على مزايا النموذج المقترح، خاصة عند مقارنته بالخوارزميات الجينية مع SVM (81.5%) وkNN (81.9%)، بالإضافة إلى M-ABC مع SVM (84.1%). تؤكد هذه النتائج على إمكانيات الإطار المقترح في تعزيز دقة التنبؤ بأمراض القلب.

المناقشة

في هذا البحث، يقدم المؤلفون منهجية جديدة للتنبؤ بأمراض القلب تدمج خوارزمية مستعمرة النحل الاصطناعية المعدلة (M-ABC) مع أقرب الجيران (k-NN) لاختيار الميزات. يتم تعزيز خوارزمية M-ABC من خلال دمج خوارزمية اليراعة، التي تحسن دقة اختيار مجموعة الميزات من خلال الاستفادة من شدة الضوء وحركة اليراعات. تهدف المنهجية إلى تحسين عملية تصنيف k-NN، مما يوضح أن الجمع بين M-ABC وk-NN يعزز بشكل كبير الأداء التنبؤي مقارنةً بالطرق التقليدية.

تم هيكلة الورقة لمراجعة الأعمال ذات الصلة أولاً حول اختيار الميزات وتقنيات التعلم الآلي لتنبؤ أمراض القلب، تليها وصف تفصيلي للمنهجية المقترحة والتقييمات التجريبية. تشير النتائج إلى أن M-ABC مع k-NN يحقق أعلى دقة بنسبة 89.7% خلال التدريب، مع قدرات تعميم قوية تتجلى في مراحل التحقق والاختبار (85.8% و84.6%، على التوالي). تشير النتائج إلى أن هذا النهج المتكامل لا يحسن فقط دقة التنبؤ ولكن أيضًا يعالج مخاوف الخصوصية من خلال السماح بتدريب النموذج محليًا على أجهزة IoMT قبل تجميع النتائج في بيئة سحابية. يقترح المؤلفون أن هذا الإطار يمكن تكييفه لتنبؤ أمراض مختلفة، مما يعزز قابليته للتطبيق في بيئات الرعاية الصحية.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-78021-1
PMID: https://pubmed.ncbi.nlm.nih.gov/39482391
Publication Date: 2024-10-31
Author(s): Muhammad Amir Khan et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

The research paper presents a novel approach for early diagnosis of heart disease using a classification model that integrates a Modified Artificial Bee Colony (M-ABC) algorithm with k-Nearest Neighbors (KNN) for optimal feature selection. The study emphasizes the importance of identifying significant attributes from the dataset to enhance classification accuracy and reduce training time. By employing the modified bee algorithm, the proposed framework retains only the most informative features during the classification-training phase, thereby facilitating more accurate predictions and informed decision-making in healthcare settings.

In the conclusion, the authors highlight the effectiveness of their proposed model, which outperforms baseline algorithms, including Genetic Algorithm with Support Vector Machine (SVM) and KNN, in various performance metrics such as accuracy, sensitivity, recall, F-score, precision, and specificity. Specifically, the model demonstrates improvements of 2.6% in accuracy and other metrics, indicating its potential for practical application in heart disease prediction while addressing privacy concerns. However, the authors acknowledge limitations, such as challenges in handling multiple Internet of Medical Things (IoMT) client sites and the sensitivity of the model to learning rate adjustments. They suggest that future research could explore advanced machine learning techniques, including deep learning models, to further enhance prediction capabilities for various serious illnesses.

Methods

The research paper presents a federated learning framework designed to predict heart disease within a privacy-aware healthcare system that utilizes the Internet of Medical Things (IoMT). This framework addresses the privacy challenges posed by regulations such as GDPR, the China Cyber Security Law, and CCPA, which complicate traditional machine learning approaches that require processing sensitive user data. The study employs the UCI Cleveland heart disease dataset, consisting of 303 records and 14 attributes, to train and evaluate the proposed model. Missing values for specific features were handled using mean imputation, while categorical features were encoded through one-hot encoding.

For experimental evaluation, the authors conducted simulations over four thousand rounds using a Python environment with PyTorch. The simulations were executed on an Intel® Core™ i7-8550 CPU at 4 GHz, ensuring consistency in the experimental setup. Data preprocessing steps included cleaning, standardization, encoding, and feature selection using MABC for dimensionality reduction. To mitigate class imbalance, the synthetic minority over-sampling technique (SMOTE) was applied. This comprehensive methodology supports the claims regarding the robustness and effectiveness of the proposed framework in predicting heart disease, as indicated by the reported accuracy, precision, recall, and F1-score metrics.

Results

The proposed framework for predicting heart disease demonstrates superior performance compared to existing techniques, including Genetic Algorithms combined with Support Vector Machines (SVM) and k-Nearest Neighbors (kNN). Specifically, the framework achieves a sensitivity of 85.6%, indicating its strong capability to accurately identify positive cases of heart disease. This sensitivity surpasses that of other evaluated methods, as illustrated in Table 3 and Figure 4.

Additionally, the framework exhibits a recall rate of 84.5%, reflecting its effectiveness in capturing all instances of favorable outcomes. This recall rate also outperforms those of the other algorithms assessed. The comparative results highlight the advantages of the proposed model, particularly when juxtaposed with Genetic Algorithms with SVM (81.5%) and kNN (81.9%), as well as M-ABC with SVM (84.1%). These findings underscore the potential of the proposed framework in enhancing heart disease prediction accuracy.

Discussion

In this research, the authors introduce a novel methodology for heart disease prediction that integrates the Modified Artificial Bee Colony (M-ABC) algorithm with k-nearest neighbors (k-NN) for feature selection. The M-ABC algorithm is enhanced by incorporating the Firefly Algorithm, which improves the precision of feature group selection by leveraging the light intensity and movement of fireflies. The methodology aims to optimize the k-NN classification process, demonstrating that the combination of M-ABC and k-NN significantly enhances predictive performance compared to traditional methods.

The paper is structured to first review related works on feature selection and machine learning techniques for heart disease prediction, followed by a detailed description of the proposed methodology and experimental evaluations. The results indicate that the M-ABC with k-NN achieves the highest accuracy of 89.7% during training, with strong generalization capabilities reflected in validation and testing phases (85.8% and 84.6%, respectively). The findings suggest that this integrated approach not only improves prediction accuracy but also addresses privacy concerns by allowing local model training on IoMT devices before aggregating results in a cloud environment. The authors propose that this framework can be adapted for predicting various diseases, thereby enhancing its applicability in healthcare settings.

شارك: