الكشف عن حالات الأمراض القلبية الوعائية باستخدام خوارزميات التعلم الآلي المتقدمة المعتمدة على الأشجار
Detection of cardiovascular disease cases using advanced tree-based machine learning algorithms

شارك:
المجلة: Scientific Reports، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41598-024-72819-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39333550
تاريخ النشر: 2024-09-27
المؤلف: Fariba Asadi وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تدرس هذه الدراسة فعالية طرق التعلم الآلي المعتمدة على الأشجار المختلفة في الكشف عن أمراض القلب والأوعية الدموية (CVD) باستخدام بيانات من 9,499 مشاركًا و38 متغيرًا. كان الهدف الرئيسي هو تحديد النموذج الذي يتمتع بأعلى أداء تنبؤي، مع التركيز على وجود CVD كمتغير مستهدف مع الأخذ في الاعتبار التجميع حسب القرية. شملت النماذج التي تم تقييمها أشجار القرار القياسية، والغابات العشوائية، وأشجار النموذج الخطي المختلط العام (GLMM)، وغابات التأثير المختلط العشوائية العامة (GMERF). أظهرت التحليلات أن خمسة متغيرات رئيسية – العمر، مستويات البروتين الدهني منخفض الكثافة (LDL)، التاريخ العائلي لأمراض القلب، النشاط البدني، وارتفاع ضغط الدم – كانت الأكثر أهمية في التنبؤ بـ CVD.

تم قياس أداء النماذج باستخدام المساحة تحت منحنى التشغيل (AUC)، حيث حصلت أشجار القرار على درجة 0.56، والغابات العشوائية على 0.73، وأشجار GLMM على 0.78، وGMERF على 0.80، مما يشير إلى أن نموذج GMERF قدم أفضل دقة تنبؤية. تشير النتائج إلى أن استخدام تقنيات التعلم الآلي التي تأخذ في الاعتبار الطبيعة المجمعة للبيانات يمكن أن يعزز القدرات التنبؤية ويعلم استراتيجيات الوقاية المستهدفة لأمراض القلب والأوعية الدموية. تؤكد هذه الدراسة على إمكانيات الأساليب المتقدمة في التعلم الآلي في تحسين النتائج الصحية المتعلقة بأمراض القلب والأوعية الدموية.

مقدمة

في هذه الدراسة، قام الباحثون بالتحقيق في 38 متنبئًا محتملاً لأمراض القلب والأوعية الدموية (CVD)، والتي تشمل مجموعة واسعة من العوامل. وشملت هذه المتغيرات الديموغرافية مثل العمر، الجنس، والحالة الاجتماعية؛ عادات نمط الحياة مثل أنماط النوم ومدخول الطاقة الغذائي؛ الحالة الاجتماعية والاقتصادية المقاسة من خلال مؤشر درجة الثروة (WSI)؛ والقياسات الأنثروبومترية بما في ذلك مؤشر كتلة الجسم (BMI) ونسبة الخصر إلى الطول. بالإضافة إلى ذلك، درست الدراسة مكافئات التمثيل الغذائي للمهام (MET)، وأنماط النظام الغذائي التي تم تقييمها من خلال مؤشر الالتهاب الغذائي (DII)، ومؤشرات حيوية دموية مختلفة (مثل LDL، HDL، الدهون الثلاثية)، والتاريخ الطبي المتعلق بحالات مثل السكري وارتفاع ضغط الدم. كما تم أخذ التاريخ الصحي العائلي، والتدخين، واستهلاك الكحول، واستخدام التبغ في الاعتبار.

كان الناتج الرئيسي الذي يهم هو وجود CVD في المرحلة الأساسية من دراسة FACS، حيث تم تصنيف المشاركين الذين تم تشخيصهم بفشل القلب (HF) أو مرض القلب الإقفاري (IHD) على أنهم يعانون من CVD. أخذ التحليل في الاعتبار تأثيرات التجميع من خلال اعتبار القرى كأثر عشوائي، مما يعزز من قوة النتائج.

طرق

يستعرض قسم “المواد والطرق” تصميم التجربة والإجراءات المستخدمة في الدراسة. يوضح المواد المحددة المستخدمة، بما في ذلك أي مواد كيميائية، معدات، وعينات بيولوجية، لضمان إمكانية تكرار البحث. تشمل المنهجية التقنيات المستخدمة لجمع البيانات وتحليلها، بما في ذلك الأساليب الإحصائية المطبقة لتفسير النتائج.

بالإضافة إلى ذلك، قد يصف القسم إعداد التجربة، الضوابط، وأي بروتوكولات ذات صلة تم اتباعها لضمان نزاهة وموثوقية النتائج. يسمح هذا النهج الشامل بفهم واضح لكيفية إجراء البحث ويدعم موثوقية الاستنتاجات المستخلصة من البيانات.

نتائج

في هذه الدراسة، تم تحليل بيانات من 9,499 مشاركًا (4,199 ذكور و5,300 إناث) بمتوسط عمر 48.95 عامًا لتحديد المتنبئين بأمراض القلب والأوعية الدموية (CVD). استخدمت عملية اختيار الميزات حزمة R ‘Boruta’، مما أدى إلى إزالة عدة متغيرات، بما في ذلك مدخول الطاقة الغذائي والتاريخ العائلي للسكري وارتفاع ضغط الدم. في النهاية، تم الاحتفاظ بـ 29 ميزة للنموذج التنبؤي النهائي، والذي شمل عوامل ديموغرافية، خيارات نمط الحياة، والتاريخ الطبي.

تم تقييم أداء نماذج مختلفة، حيث أظهر نموذج الغابة العشوائية ذات التأثيرات المختلطة العامة (gmerf) أعلى قوة تنبؤية (AUC = 0.80)، تلاه عن كثب نموذج شجرة GLMM. أظهر الأخير أعلى دقة (0.81) ولكنه كان لديه خصوصية أقل (0.59). تشمل المتنبئات الرئيسية لـ CVD التي تم تحديدها عبر النماذج العمر، مستويات الكوليسترول LDL، التاريخ العائلي لأمراض القلب، مستوى النشاط البدني، وارتفاع ضغط الدم. بالمقابل، اعتبرت متغيرات مثل استخدام الكحول ونسبة الخصر إلى الطول أقل أهمية وتم استبعادها من نموذج الشجرة النهائي. تشير نتائج الدراسة إلى أن خصائص معينة يمكن أن تصنف بفعالية الأفراد المعرضين لخطر CVD، مع توفير العقد النهائية رؤى حول احتمال وجود المرض بناءً على المعايير المحددة.

مناقشة

في هذه الدراسة، استخدم المؤلفون بيانات من دراسة فاسا التعاونية (FACS) لتطوير نماذج التعلم الآلي التي تهدف إلى التنبؤ بأمراض القلب والأوعية الدموية (CVD) بين السكان الريفيين في إيران. شملت العينة 9,499 مشاركًا بعد معالجة البيانات المفقودة، وتم استخدام تقنيات تعلم آلي متنوعة، بما في ذلك الغابات العشوائية ذات التأثيرات المختلطة العامة (gmerf) والنماذج الخطية المختلطة العامة (GLMM). أظهر نموذج gmerf أداءً تنبؤيًا متفوقًا، حيث حقق أعلى مساحة تحت منحنى التشغيل (ROC) (AUC = 0.80)، بينما أظهر نموذج شجرة القرار أدنى خصوصية وAUC (0.55)، مما يشير إلى مشاكل محتملة في التصنيف.

شملت المتنبئات الرئيسية التي تم تحديدها العمر، مستويات الكوليسترول LDL، التاريخ العائلي لأمراض القلب، النشاط البدني، وارتفاع ضغط الدم. تتماشى هذه النتائج مع الأدبيات الموجودة، مما يعزز أهمية هذه المتغيرات في تقييم خطر CVD. تؤكد الدراسة على أهمية نهج شامل لاختيار الميزات وتطوير النماذج، داعية إلى دمج تقنيات التعلم الآلي في الممارسة السريرية لتعزيز الكشف المبكر وإدارة CVD، بهدف تحسين نتائج الصحة العامة وتقليل تكاليف الرعاية الصحية.

Journal: Scientific Reports, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41598-024-72819-9
PMID: https://pubmed.ncbi.nlm.nih.gov/39333550
Publication Date: 2024-09-27
Author(s): Fariba Asadi et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

This study investigates the efficacy of various tree-based machine learning methods for detecting cardiovascular disease (CVD) using data from 9,499 participants and 38 variables. The primary objective was to identify the model with the highest predictive performance, focusing on the presence of CVD as the target variable while accounting for clustering by village. The models evaluated included standard decision trees, random forests, Generalized Linear Mixed Model trees (GLMM trees), and Generalized Mixed Effect Random Forests (GMERF). The analysis revealed that five key variables—age, low-density lipoprotein (LDL) levels, family history of cardiac disease, physical activity, and hypertension—were the most significant predictors of CVD.

The performance of the models was quantified using the area under the ROC curve (AUC), yielding scores of 0.56 for the decision tree, 0.73 for random forests, 0.78 for GLMM trees, and 0.80 for GMERF, indicating that the GMERF model provided the best predictive accuracy. The findings suggest that employing machine learning techniques that consider the clustered nature of the data can enhance predictive capabilities and inform targeted prevention strategies for CVD. This research underscores the potential of advanced machine learning approaches in improving health outcomes related to cardiovascular diseases.

Introduction

In this study, the researchers investigated 38 potential predictors of cardiovascular disease (CVD), encompassing a wide range of factors. These included demographic variables such as age, gender, and marital status; lifestyle habits like sleep patterns and dietary energy intake; socio-economic status measured by a wealth score index (WSI); and anthropometric measurements including body mass index (BMI) and waist-to-height ratio. Additionally, the study examined metabolic equivalents of task (MET), dietary patterns assessed through the Dietary Inflammatory Index (DII), various blood biomarkers (e.g., LDL, HDL, triglycerides), and medical history related to conditions such as diabetes and hypertension. Family health history, smoking, alcohol consumption, and tobacco use were also considered.

The primary outcome of interest was the presence of CVD at the baseline phase of the FACS study, with participants diagnosed with heart failure (HF) or ischemic heart disease (IHD) classified as having CVD. The analysis accounted for clustering effects by treating villages as a random effect, thereby enhancing the robustness of the findings.

Methods

The “Materials and Methods” section outlines the experimental design and procedures employed in the study. It details the specific materials used, including any reagents, equipment, and biological samples, ensuring reproducibility of the research. The methodology encompasses the techniques for data collection and analysis, including statistical methods applied to interpret the results.

Additionally, the section may describe the experimental setup, controls, and any relevant protocols followed to ensure the integrity and validity of the findings. This comprehensive approach allows for a clear understanding of how the research was conducted and supports the reliability of the conclusions drawn from the data.

Results

In this study, data from 9,499 participants (4,199 males and 5,300 females) with a mean age of 48.95 years were analyzed to identify predictors of cardiovascular disease (CVD). The feature selection process utilized the ‘Boruta’ R package, resulting in the removal of several variables, including dietary energy intake and family history of diabetes and hypertension. Ultimately, 29 features were retained for the final predictive model, which included demographic factors, lifestyle choices, and medical history.

The performance of various models was evaluated, with the generalized mixed-effects random forest (gmerf) model demonstrating the highest predictive power (AUC = 0.80), followed closely by the glmm tree model. The latter exhibited the highest accuracy (0.81) but had lower specificity (0.59). Key predictors of CVD identified across models included age, LDL cholesterol levels, family history of cardiac disease, physical activity level, and hypertension. Conversely, variables such as alcohol use and waist-to-height ratio were deemed less important and were excluded from the final tree model. The study’s findings suggest that specific characteristics can effectively classify individuals at risk for CVD, with terminal nodes providing insights into the likelihood of disease presence based on identified criteria.

Discussion

In this study, the authors utilized data from the Fasa Cohort Study (FACS) to develop machine learning models aimed at predicting cardiovascular disease (CVD) among rural populations in Iran. The cohort included 9,499 participants after addressing missing data, and various machine learning techniques were employed, including generalized mixed-effect random forests (gmerf) and generalized linear mixed models (GLMM). The gmerf model demonstrated superior predictive performance, achieving the highest area under the receiver operating characteristic (ROC) curve (AUC = 0.80), while the decision tree model exhibited the lowest specificity and AUC (0.55), indicating potential misclassification issues.

Key predictors identified included age, low-density lipoprotein (LDL) cholesterol levels, family history of cardiac disease, physical activity, and hypertension. These findings align with existing literature, reinforcing the significance of these variables in CVD risk assessment. The study emphasizes the importance of a comprehensive approach to feature selection and model development, advocating for the integration of machine learning techniques in clinical practice to enhance early detection and management of CVD, ultimately aiming to improve public health outcomes and reduce healthcare costs.

شارك: