التنبؤ بمرض السكري لدى البالغين: تحديد الميزات المهمة في البيانات غير المتوازنة على مدى دراسة طولية لمدة 5 سنوات باستخدام خوارزمية التعلم الآلي
Predicting diabetes in adults: identifying important features in unbalanced data over a 5-year cohort study using machine learning algorithm

شارك:
المجلة: BMC Medical Research Methodology، المجلد: 24، العدد: 1
DOI: https://doi.org/10.1186/s12874-024-02341-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39333899
تاريخ النشر: 2024-09-27
المؤلف: Maryam Talebi Moghaddam وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تستكشف هذه الدراسة التحديات التي تطرحها مجموعات البيانات غير المتوازنة في النمذجة التنبؤية، وخاصة في سياق توقع مرض السكري باستخدام تقنيات التعلم الآلي. باستخدام بيانات من دراسة فاسا للبالغين (FACS) التي تشمل 10,000 مشارك على مدى 5 سنوات، تستخدم البحث تدخلات مختلفة على مستوى البيانات وعلى مستوى الخوارزميات، بما في ذلك SMOTE وADASYN وRandom Over Sampling، بالتزامن مع مصنفات مثل Random Forest وGradient Boosting وDecision Tree وMulti-Layer Perceptron (MLP). يتم تقييم أداء هذه النماذج باستخدام مقاييس مثل درجة F1 وAUC وG-means لضمان تقييم شامل لدقة النموذج وتمييزه، خاصة في السيناريوهات غير المتوازنة.

تكشف النتائج عن مؤشرات خطر السكري المحددة حسب الجنس، حيث تعتبر الدهون الثلاثية (TG) ومعدل الأيض الأساسي (BMR) والكوليسترول الكلي (CHOL) الأكثر تأثيرًا لدى الإناث، بينما يعتبر مؤشر كتلة الجسم (BMI) وإنزيم ناقل الأمين الجلوتامات (SGOT) وغلوتاميل غاما (GGT) أساسية للذكور. ومن الجدير بالذكر أن BMI تم تحديده كأهم متغير عبر مجموعة البيانات بأكملها. تسلط الدراسة الضوء على أن الجمع بين ADASYN ومصنف MLP حقق أعلى مقاييس أداء، حيث حقق درجة F1 قدرها 82.17 ± 3.38 وAUC قدرها 89.61 ± 2.09. تؤكد النتائج فعالية تقنيات إعادة أخذ العينات المتقدمة في تعزيز الدقة التنبؤية، مما يبرز الحاجة إلى استراتيجيات محددة حسب الجنس في الوقاية من السكري وإدارته.

في الختام، تُظهر هذه البحث إمكانية نماذج التعلم الآلي المدعومة بتقنيات توازن البيانات لتحسين توقع مرض السكري. ستركز الأعمال المستقبلية على تحسين طرق اختيار الميزات واستكشاف أساليب خوارزمية متقدمة، مثل طرق التجميع الهجينة والتعلم المعزز، لتعزيز أداء النموذج بشكل أكبر ومعالجة عدم التوازن في البيانات بشكل تكيفي. هذه الرؤى ضرورية لتطوير استراتيجيات وقائية وإدارية مستهدفة مصممة لتناسب ملفات المخاطر المحددة حسب الجنس في مرض السكري.

مقدمة

تتناول مقدمة هذه الورقة البحثية الأزمة الصحية العالمية المتزايدة التي يسببها مرض السكري من النوع 2، وهو اضطراب استقلابي مزمن يتميز بمقاومة الأنسولين أو إنتاج الأنسولين غير الكافي. يرتبط هذا الحالة بمضاعفات خطيرة، بما في ذلك الأمراض القلبية الوعائية، وفشل الكلى، وزيادة معدلات الوفيات، حيث تتوقع الاتحاد الدولي للسكري ارتفاع عدد المصابين بالسكري في العالم من 382 مليون في عام 2013 إلى 592 مليون بحلول عام 2035. من المتوقع أيضًا أن ترتفع نسبة انتشار المرض بين البالغين من 6.4% في عام 2010 إلى 7.7% بحلول عام 2030، مما يبرز الحاجة الملحة لأدوات تنبؤية فعالة لإدارة تأثير المرض.

استجابةً لهذه الحاجة، تؤكد الدراسة على تطبيق تقنيات التعلم الآلي (ML) لتوقع ظهور مرض السكري ومضاعفاته. لقد أظهرت نماذج ML المختلفة، مثل Random Forest وK-NN والشبكات العصبية، فعاليتها في توقع المخاطر وإدارتها. ومع ذلك، تنشأ تحديات كبيرة من عدم توازن البيانات في مجموعات البيانات السريرية، مما قد يؤثر على دقة النماذج التنبؤية. لمعالجة هذه المشكلة، تنفذ الدراسة تدخلات على مستوى البيانات وعلى مستوى الخوارزميات، بما في ذلك زيادة العينات، وتقليل العينات، وتقنيات أخذ العينات الهجينة، مع تقييم أداء خوارزميات ML المختلفة باستخدام مقاييس مثل درجة F1 وAUC ومؤشرات G-means لتحديد أكثر الأساليب التنبؤية فعالية لحدوث مرض السكري من النوع 2.

طرق

في هذا القسم، يحدد المؤلفون الإطار المنهجي المستخدم في دراستهم، والذي يشمل جمع البيانات، والمعالجة المسبقة، والتعامل مع عدم توازن البيانات، مما يؤدي إلى اختيار المصنفات ومقاييس التقييم. يتم تقديم تمثيل بصري (الشكل 1) لإرشاد القراء خلال سير العمل البحثي، مع التأكيد على أهمية كل خطوة في تحليل اكتشاف مرض السكري.

لمعالجة عدم توازن البيانات، يناقش المؤلفون طرق أخذ العينات المختلفة التي تهدف إلى ضمان التعلم العادل من كلا الفئتين. تشمل تقنيات زيادة العينات الرئيسية زيادة العينات العشوائية، التي تكرر أمثلة الفئة الأقل، وطرق أكثر تعقيدًا مثل SMOTE (تقنية زيادة العينات الأقلية الاصطناعية) وADASYN (نهج أخذ العينات الاصطناعية التكيفية) وBorderline SMOTE وKMeans SMOTE وSmotified GAN، وكلها مصممة لتحسين أداء النموذج من خلال توليد عينات اصطناعية مع تقليل مخاطر الإفراط في التكيف. بالإضافة إلى ذلك، يتم تقديم طرق أخذ العينات الهجينة، مثل SMOTEENN وSMOTETomek، التي تجمع بين استراتيجيات زيادة العينات وتقليل العينات لتحقيق توازن فعال في مجموعة البيانات مع تقليل فقد المعلومات وتوضيح حدود القرار. تعتبر هذه المنهجيات حاسمة لتحسين دقة وحساسية نماذج التصنيف في اكتشاف مرض السكري.

مناقشة

تسلط قسم المناقشة في الورقة البحثية الضوء على التقدم في تقنيات التعلم الآلي التي تهدف إلى معالجة التحديات التي تطرحها مجموعات البيانات غير المتوازنة، وخاصة في توقع مرض السكري. لقد أظهرت الدراسات الحديثة فعالية طرق إعادة أخذ العينات المختلفة، مثل زيادة العينات العشوائية وSMOTE وADASYN، في تعزيز الأداء التنبؤي عبر مجموعات بيانات صحية متعددة. ومن الجدير بالذكر أن العمل الذي قام به O. Olawale Awe وآخرون (2024) حدد طريقة أخذ العينات الأقرب المتكررة المدمجة مع الانحدار اللوجستي كوسيلة لتحقيق تحسينات كبيرة في دقة التنبؤ. وقد أكدت دراسات أخرى، بما في ذلك تلك التي أجراها Wahyu Nugraha وآخرون (2023) وHirani Hairani وDadang Priyanto (2023)، فعالية مجموعات معينة من تقنيات إعادة أخذ العينات والمصنفات، مثل SMOTE مع أشجار القرار وSMOTE-ENN مع Random Forest، على التوالي.

شملت منهجية الدراسة جمع بيانات شاملة من دراسة فاسا للبالغين، مع التركيز على حدوث مرض السكري من النوع 2 على مدى 5 سنوات. تم إنشاء إطار معالجة مسبقة قوي لضمان سلامة البيانات، بما في ذلك التعامل مع القيم المفقودة وتطبيع مجموعة البيانات. كشفت التحليلات أن نموذج Multi-Layer Perceptron (MLP)، خاصة عند اقترانه بزيادة العينات العشوائية، حقق أعلى مقاييس أداء، بما في ذلك درجة F1 قدرها 82.97 وAUC قدرها 89.25. تؤكد النتائج على الدور الحاسم لتقنيات توازن البيانات في تحسين موثوقية التنبؤ، مما يشير إلى أن الأساليب المخصصة لإعادة أخذ العينات واختيار المصنفات يمكن أن تعزز بشكل كبير أداء النموذج في التطبيقات الطبية. تشمل اتجاهات البحث المستقبلية استكشاف طرق اختيار الميزات ودمج استراتيجيات التعلم الآلي المتقدمة لتحسين القدرات التنبؤية لإدارة مرض السكري.

Journal: BMC Medical Research Methodology, Volume: 24, Issue: 1
DOI: https://doi.org/10.1186/s12874-024-02341-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39333899
Publication Date: 2024-09-27
Author(s): Maryam Talebi Moghaddam et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

This study investigates the challenges posed by imbalanced datasets in predictive modeling, specifically in the context of diabetes prediction using machine learning techniques. Utilizing data from the Fasa Adult Cohort Study (FACS) involving 10,000 participants over a 5-year follow-up, the research employs various data-level and algorithm-level interventions, including SMOTE, ADASYN, and Random Over Sampling, in conjunction with classifiers such as Random Forest, Gradient Boosting, Decision Tree, and Multi-Layer Perceptron (MLP). The performance of these models is assessed using metrics like F1 score, AUC, and G-means to ensure a comprehensive evaluation of model accuracy and discrimination, particularly in imbalanced scenarios.

The findings reveal significant gender-specific predictors of diabetes risk, with triglyceride (TG), basal metabolic rate (BMR), and total cholesterol (CHOL) being the most influential for females, while body mass index (BMI), serum glutamate oxaloacetate transaminase (SGOT), and gamma-glutamyl (GGT) are key for males. Notably, BMI is identified as the most important variable across the entire dataset. The study highlights that the combination of ADASYN with the MLP classifier yielded the highest performance metrics, achieving an F1 score of 82.17 ± 3.38 and an AUC of 89.61 ± 2.09. The results underscore the effectiveness of advanced resampling techniques in enhancing predictive accuracy, emphasizing the need for gender-specific strategies in diabetes prevention and management.

In conclusion, this research demonstrates the potential of machine learning models augmented with data balancing techniques to improve diabetes prediction. Future work will focus on refining feature selection methods and exploring advanced algorithmic approaches, such as hybrid ensemble methods and reinforcement learning, to further enhance model performance and adaptively address data imbalance. These insights are crucial for developing targeted prevention and management strategies tailored to gender-specific risk profiles in diabetes.

Introduction

The introduction of this research paper addresses the growing global health crisis posed by Type 2 diabetes, a chronic metabolic disorder marked by insulin resistance or inadequate insulin production. This condition is linked to severe complications, including cardiovascular diseases, kidney failure, and increased mortality rates, with the International Diabetes Federation projecting a rise in the global diabetes population from 382 million in 2013 to 592 million by 2035. The prevalence among adults is also expected to increase from 6.4% in 2010 to 7.7% by 2030, highlighting the urgent need for effective predictive tools to manage the disease’s impact.

In response to this need, the study emphasizes the application of machine learning (ML) techniques for predicting the onset of diabetes and its complications. Various ML models, such as Random Forest, K-NN, and neural networks, have demonstrated effectiveness in risk prediction and management. However, a significant challenge arises from data imbalance in clinical datasets, which can compromise the accuracy of predictive models. To address this issue, the study implements both data-level and algorithm-level interventions, including oversampling, undersampling, and hybrid sampling techniques, while evaluating the performance of different ML algorithms using metrics like the F1 score, AUC, and G-means indices to identify the most effective predictive approaches for Type 2 diabetes incidence.

Methods

In this section, the authors outline the methodological framework employed in their study, which encompasses data collection, preprocessing, and the handling of data imbalance, culminating in classifier selection and evaluation metrics. A visual representation (Fig. 1) is provided to guide readers through the research workflow, emphasizing the importance of each step in the analysis of diabetes detection.

To address data imbalance, the authors discuss various sampling methods aimed at ensuring equitable learning from both classes. Key oversampling techniques include Random Oversampling, which duplicates minority class examples, and more sophisticated methods like SMOTE (Synthetic Minority Oversampling Technique), ADASYN (Adaptive Synthetic Sampling Approach), Borderline SMOTE, KMeans SMOTE, and Smotified GAN, each designed to enhance model performance by generating synthetic samples while mitigating overfitting risks. Additionally, hybrid sampling methods, such as SMOTEENN and SMOTETomek, are introduced, which combine oversampling and undersampling strategies to balance the dataset effectively while minimizing information loss and clarifying decision boundaries. These methodologies are critical for improving the accuracy and sensitivity of classification models in diabetes detection.

Discussion

The discussion section of the research paper highlights the advancements in machine learning techniques aimed at addressing the challenges posed by imbalanced datasets, particularly in predicting diabetes. Recent studies have demonstrated the effectiveness of various resampling methods, such as Random Oversampling, SMOTE, and ADASYN, in enhancing predictive performance across multiple health-related datasets. Notably, the work by O. Olawale Awe et al. (2024) identified the Repeated Nearest Neighbor Sampling method combined with logistic regression as yielding significant improvements in prediction accuracy. Other studies, including those by Wahyu Nugraha et al. (2023) and Hirani Hairani and Dadang Priyanto (2023), further corroborated the efficacy of specific combinations of resampling techniques and classifiers, such as SMOTE with decision trees and SMOTE-ENN with Random Forest, respectively.

The study’s methodology involved a comprehensive data collection from the Fasa Adult Cohort Study, focusing on the 5-year cumulative incidence of type 2 diabetes. A robust preprocessing framework was established to ensure data integrity, including handling missing values and normalizing the dataset. The analysis revealed that the Multi-Layer Perceptron (MLP) model, particularly when paired with Random Over Sampling, achieved the highest performance metrics, including an F1 score of 82.97 and an AUC of 89.25. The findings emphasize the critical role of data balancing techniques in improving predictive reliability, suggesting that tailored approaches to resampling and classifier selection can significantly enhance model performance in medical applications. Future research directions include exploring feature selection methods and integrating advanced machine learning strategies to further refine predictive capabilities for diabetes management.

شارك: