قابلية تعميم نماذج التعلم الآلي لاكتشاف السكري: دراسة باستخدام بيانات زراعة جزر النوردية وبيانات PIMA
Generalizability of machine learning models for diabetes detection a study with nordic islet transplant and PIMA datasets

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-87471-0
PMID: https://pubmed.ncbi.nlm.nih.gov/39915538
تاريخ النشر: 2025-02-06
المؤلف: Dinesh Chellappan وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تبحث الدراسة في تطبيق تقنيات التعلم الآلي للكشف المبكر عن داء السكري (DM) باستخدام بيانات التعبير الجيني من الميكروأري ومجموعة بيانات السكري الهندية PIMA (PIDD). تم استخدام طريقة استخراج ميزات هجينة تجمع بين خوارزمية مستعمرة النحل الاصطناعية (ABC) وتحسين سرب الجسيمات (PSO)، تلتها خوارزميات اختيار الميزات الميتاهيرستية مثل البحث الهارموني (HS) وخوارزمية اليعسوب (DFA) وخوارزمية رعي الفيلة (EHA). تم تقييم أداء مصنفات مختلفة، بما في ذلك الانحدار غير الخطي (NLR) والانحدار الخطي (LR) ونموذج المزيج الغاوسي (GMM) والتوقع الأقصى (EM) وتحليل التمييز الخطي بايزي (BLDA) ومصنف التمييز Softmax (SDC) وآلة الدعم الناقل مع نواة دالة القاعدة الشعاعية (SVM-RBF)، على مجموعتين من البيانات: برنامج زراعة الجزر النوردي (NITP) وPIDD.

تشير النتائج إلى أن النهج الهجين يعزز بشكل كبير دقة التصنيف، حيث حقق استخراج الميزات ABC-PSO مع اختيار EHA أعلى دقة بنسبة 97.14% على NITP و98.13% على PIDD. تتجاوز هذه الدقة تلك التي تم الحصول عليها باستخدام تقنيات فردية، مما يظهر قوة الطريقة المقترحة عبر مجموعات بيانات مختلفة. تستنتج الدراسة أن طرق استخراج الميزات واختيارها الهجينة تحدد بشكل فعال الميزات الأكثر معلوماتية لتوقع السكري، مما يساهم في التقدم في تطبيقات التعلم الآلي في تشخيص السكري.

النتائج

تقدم قسم النتائج في الدراسة تقييمًا شاملاً لنماذج التصنيف المختلفة لتحديد مرضى السكري، باستخدام التحقق المتقاطع بعشر طيات لضمان القوة وتقليل الإفراط في التكيف. تم حساب مقاييس الأداء، بما في ذلك الدقة، ودرجة F1، ومعامل ارتباط ماثيو (MCC)، ومقياس جاكارد، ومعدل الخطأ، وإحصائية كابا، باستخدام مصفوفة الارتباك، مما يوفر تقييمًا شاملاً لقدرات النماذج. من الجدير بالذكر أن آلة الدعم الناقل مع دالة القاعدة الشعاعية (SVM(RBF)) أظهرت أداءً متفوقًا عبر كل من بيانات الجينات الميكروأري ومجموعة بيانات السكري الهندية PIMA، محققة دقة بنسبة 88.57% ودرجة F1 بنسبة 81.82% على بيانات الميكروأري، مع الحفاظ على نطاق دقة يتراوح بين 88.31-89.74% على مجموعة بيانات PIMA.

تسلط التحليلات الضوء أيضًا على فعالية تقنيات استخراج الميزات، وخاصة طريقة البحث الهارموني، التي عززت بشكل كبير أداء SVM(RBF) على بيانات الميكروأري، محققة دقة تتجاوز 91%. بينما حسنت طرق اختيار الميزات مثل البحث الهارموني وتحليل الميزات التفاضلية (DFA) والبحث الهارموني المعزز (EHA) نتائج التصنيف، كان تأثيرها متغيرًا مع أبعاد البيانات. تميزت EHA بشكل خاص في السياقات عالية الأبعاد، بينما حافظت SVM(RBF) على أداء قوي عبر تقنيات اختيار الميزات المختلفة في مجموعات بيانات منخفضة الأبعاد مثل PIMA. بالإضافة إلى ذلك، أشار تحليل مخطط التشتت إلى وجود ارتباط إيجابي ضعيف بين معدل الإيجابيات الكاذبة (FPR) والقيمة المتوقعة السلبية (NPV)، مما يشير إلى أنه مع زيادة FPR، قد ترتفع NPV أيضًا، وإن لم يكن بطريقة خطية صارمة. تؤكد هذه النتيجة على الحاجة إلى مزيد من الاستكشاف للعلاقة بين هذه المقاييس في سياق تصنيف السكري.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على التحديات والمنهجيات المستخدمة في استخدام بيانات التعبير الجيني لتشخيص السكري. تؤكد على أهمية تقنيات استخراج الميزات واختيارها، وخاصة خوارزمية مستعمرة النحل الاصطناعية (ABC) ونموذجها الهجين مع تحسين سرب الجسيمات (ABC-PSO)، في معالجة الأبعاد العالية وتعقيد مجموعات بيانات الميكروأري. تهدف هذه التقنيات إلى تحديد أكثر مجموعات الجينات معلوماتية، مما يعزز أداء مصنفات مختلفة مثل الانحدار غير الخطي (NLR) والانحدار اللوجستي (LR) وآلات الدعم الناقل (SVM-RBF). توضح الدراسة فعالية هذه الطرق من خلال التقييمات على مجموعتين من البيانات المتاحة للجمهور، زراعة النوردي وPIMA، باستخدام مقاييس مثل المساحة تحت المنحنى (AUC) ودرجة F1.

تشير النتائج إلى أنه بينما تقلل تقنيات استخراج الميزات الأبعاد، فإن اختيار الميزات الإضافي ضروري لتحسين دقة التصنيف. تناقش الورقة تطبيق الخوارزميات الميتاهيرستية، بما في ذلك البحث الهارموني (HS) وخوارزمية اليعسوب (DFA) وخوارزمية رعي الفيلة (EHA)، لتنقيح اختيار الميزات الحرجة. تكشف التحليلات الإحصائية، بما في ذلك اختبارات t، أن ABC وABC-PSO تتفوقان بشكل كبير على الاختيار العشوائي في تحديد الميزات ذات الصلة، بينما تظهر HS وDFA وEHA فعالية متفاوتة. بشكل عام، تؤكد الأبحاث على إمكانية دمج تقنيات الحوسبة المتقدمة مع البيانات البيولوجية لتعزيز الكشف عن السكري، مع الإشارة أيضًا إلى التوازن بين قابلية تفسير النموذج والأداء.

القيود

تقدم الدراسة عدة قيود تبرز مجالات البحث المستقبلي. أولاً، المجموعات البيانات المستخدمة، على الرغم من تنوعها، محدودة في الحجم والتمثيل الديموغرافي، مما قد يؤثر على قابلية تعميم النتائج. لتعزيز قوة النتائج، يجب أن تهدف الدراسات اللاحقة إلى دمج مجموعات بيانات أكبر وأكثر تنوعًا. بالإضافة إلى ذلك، تواجه المنهجية الهجينة المستخدمة، التي تجمع بين الكفاءة الحاسوبية والأداء التنبؤي، تحديات بسبب التكاليف الحاسوبية العالية المرتبطة بالطبيعة التكرارية لخوارزمية تحسين سرب الجسيمات بايزي التقريبية (ABC-PSO) وغيرها من الخوارزميات الميتاهيرستية. يجب أن تركز الأعمال المستقبلية على تحسين كفاءة الخوارزميات من خلال تقنيات مثل التوازي واختيار الميزات الخفيفة.

علاوة على ذلك، يفتقر النموذج الحالي إلى قابلية التفسير، وهو جانب حاسم لاتخاذ القرارات السريرية. يجب أن تستكشف الأبحاث المستقبلية طرقًا لتعزيز قابلية تفسير النموذج، مما يسهل اعتماده في البيئات السريرية. أخيرًا، تعتبر الاستراتيجيات العملية لدمج نماذج التعلم الآلي في سير العمل السريري، إلى جانب معالجة الاعتبارات الأخلاقية المتعلقة باستخدام البيانات الجينية، ضرورية للتطبيق المسؤول لهذه الأبحاث في السيناريوهات الواقعية. تشير النتائج من تحليل مصنفات مختلفة وطرق استخراج الميزات إلى معدلات دقة واعدة، مما يوحي بأن المزيد من التحسين قد يؤدي إلى تحسينات كبيرة في الأداء التنبؤي.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-87471-0
PMID: https://pubmed.ncbi.nlm.nih.gov/39915538
Publication Date: 2025-02-06
Author(s): Dinesh Chellappan et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

The research investigates the application of machine learning techniques for the early detection of Diabetes Mellitus (DM) using microarray gene expression data and the PIMA Indian Diabetes Dataset (PIDD). A hybrid feature extraction method combining Artificial Bee Colony (ABC) and Particle Swarm Optimization (PSO) was employed, followed by metaheuristic feature selection algorithms such as Harmonic Search (HS), Dragonfly Algorithm (DFA), and Elephant Herding Algorithm (EHA). The performance of various classifiers, including Non-Linear Regression (NLR), Linear Regression (LR), Gaussian Mixture Model (GMM), Expectation Maximization (EM), Bayesian Linear Discriminant Analysis (BLDA), Softmax Discriminant Classifier (SDC), and Support Vector Machine with Radial Basis Function kernel (SVM-RBF), was evaluated on two datasets: the Nordic Islet Transplant Program (NITP) and PIDD.

The results indicate that the hybrid approach significantly enhances classification accuracy, with the ABC-PSO feature extraction combined with EHA selection achieving the highest accuracy of 97.14% on the NITP and 98.13% on the PIDD. These accuracies surpass those obtained using individual techniques, demonstrating the robustness of the proposed method across different datasets. The study concludes that the hybrid feature extraction and selection methods effectively identify the most informative features for diabetes prediction, thereby contributing to advancements in machine learning applications in diabetes diagnostics.

Results

The results section of the study presents a thorough evaluation of various classification models for diabetic patient identification, employing tenfold cross-validation to ensure robustness and mitigate overfitting. Performance metrics, including accuracy, F1 score, Matthews Correlation Coefficient (MCC), Jaccard Metric, Error Rate, and Kappa statistic, were calculated using a confusion matrix, providing a comprehensive assessment of the models’ capabilities. Notably, the Support Vector Machine with Radial Basis Function (SVM(RBF)) classifier demonstrated superior performance across both microarray gene data and the PIMA Indians Diabetes Dataset, achieving an accuracy of 88.57% and an F1 score of 81.82% on the microarray data, and maintaining an accuracy range of 88.31-89.74% on the PIMA dataset.

The analysis further highlights the effectiveness of feature extraction techniques, particularly the Harmonic Search method, which significantly enhanced the performance of SVM(RBF) on microarray data, achieving accuracy exceeding 91%. While feature selection methods such as Harmonic Search, Differential Feature Analysis (DFA), and Enhanced Harmony Search (EHA) improved classification outcomes, their impact varied with data dimensionality. EHA notably excelled in high-dimensional contexts, while SVM(RBF) maintained strong performance across various feature selection techniques in lower-dimensional datasets like PIMA. Additionally, a scatter plot analysis indicated a weak positive correlation between the False Positive Rate (FPR) and Negative Predicted Value (NPV), suggesting that as FPR increases, NPV may also rise, albeit not in a strictly linear manner. This finding underscores the need for further exploration of the relationship between these metrics in the context of diabetic classification.

Discussion

The discussion section of the research paper highlights the challenges and methodologies employed in utilizing gene expression data for diabetes diagnosis. It emphasizes the importance of feature extraction and selection techniques, particularly the Artificial Bee Colony (ABC) algorithm and its hybrid variant with Particle Swarm Optimization (ABC-PSO), in addressing the high dimensionality and complexity of microarray datasets. These techniques aim to identify the most informative gene subsets, thereby enhancing the performance of various classifiers such as Non-Linear Regression (NLR), Logistic Regression (LR), and Support Vector Machines (SVM-RBF). The study demonstrates the effectiveness of these methods through evaluations on two publicly available datasets, the Nordic transplant and PIMA, using metrics like Area Under the Curve (AUC) and F1-Score.

The findings indicate that while feature extraction techniques successfully reduce dimensionality, further feature selection is necessary to improve classification accuracy. The paper discusses the application of metaheuristic algorithms, including Harmony Search (HS), Dragon Fly Algorithm (DFA), and Elephant Herding Algorithm (EHA), to refine the selection of critical features. Statistical analyses, including t-tests, reveal that ABC and ABC-PSO significantly outperform random selection in identifying relevant features, while HS, DFA, and EHA show varying effectiveness. Overall, the research underscores the potential of combining advanced computational techniques with biological data to enhance diabetes detection, while also noting the trade-off between model interpretability and performance.

Limitations

The study presents several limitations that highlight areas for future research. Firstly, the datasets utilized, while diverse, are constrained in size and demographic representation, potentially affecting the generalizability of the results. To enhance the robustness of the findings, subsequent studies should aim to incorporate larger and more varied datasets. Additionally, the hybrid methodology employed, which combines computational efficiency with predictive performance, faces challenges due to the high computational costs associated with the iterative nature of Approximate Bayesian Computation Particle Swarm Optimization (ABC-PSO) and other metaheuristic algorithms. Future work should focus on improving algorithmic efficiency through techniques such as parallelization and lightweight feature selection.

Moreover, the current model lacks interpretability, a critical aspect for clinical decision-making. Future research should explore methods to enhance model explainability, thereby facilitating its adoption in clinical settings. Finally, practical strategies for integrating machine learning models into clinical workflows, along with addressing ethical considerations related to the use of genetic data, are essential for the responsible application of this research in real-world scenarios. The findings from the analysis of various classifiers and feature extraction methods indicate promising accuracy rates, suggesting that further refinement could yield significant improvements in predictive performance.

شارك: