تحسين توقع مخاطر الأمراض القلبية الوعائية باستخدام نماذج تعلم الآلة المتكاملة: دراسة من مجموعة بيانات UK Biobank
Enhanced cardiovascular disease risk prediction using integrated machine learning models: a study from the UK Biobank cohort

شارك:
المجلة: Open Heart، المجلد: 13، العدد: 1
DOI: https://doi.org/10.1136/openhrt-2025-003827
PMID: https://pubmed.ncbi.nlm.nih.gov/41720602
تاريخ النشر: 2026-01-01
المؤلف: Qingmiao Yang وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية

نظرة عامة

تتناول هذه الورقة البحثية التحدي المستمر في التنبؤ بدقة بأمراض القلب والأوعية الدموية (CVD)، التي تُعتبر السبب الرئيسي للوفيات العالمية. غالبًا ما تؤدي نماذج التنبؤ بالمخاطر التقليدية، مثل درجة مخاطر القلب والأوعية الدموية في فرامينغهام، أداءً ضعيفًا بسبب اعتمادها على متغيرات محدودة وطرق إحصائية كلاسيكية. للتغلب على هذه القيود، تستخدم الدراسة تقنيات التعلم الآلي (ML) لتطوير نموذج شامل للتنبؤ بمخاطر CVD باستخدام بيانات من مجموعة بيانات UK Biobank، والتي تشمل الدرجات الجينية، والمعايير السريرية، وعوامل نمط الحياة. شملت التحليلات 240,644 مشاركًا واستخدمت خوارزميات ML متنوعة، محققةً منطقة تحت المنحنى (AUC) تبلغ 0.85 للنموذج المتكامل، متفوقة بشكل كبير على النماذج المستندة إلى مجالات بيانات فردية.

تشير النتائج الرئيسية إلى أن النموذج المتكامل لا يعزز فقط دقة التنبؤ ولكن أيضًا يحدد المتنبئين الرئيسيين مثل العمر، وضغط الدم الانقباضي، ومستويات السيستاتين C، ومحيط الخصر. تسلط الدراسة الضوء على تحسن قدره 0.12 في AUC مقارنةً بدرجة فرامينغهام التقليدية، مما يبرز إمكانية استخدام النموذج في التطبيقات السريرية في تحديد المخاطر المبكرة والتدخلات الشخصية. بالإضافة إلى ذلك، يحتفظ إصدار مبسط من النموذج، الذي يستخدم فقط 20 متنبئًا يمكن الحصول عليها بسهولة، بأداء تنبؤي عالٍ، مما يجعله عمليًا للاستخدام السريري. يدعو المؤلفون إلى دمج أدوات التنبؤ بالمخاطر المستندة إلى ML في استراتيجيات الوقاية من أمراض القلب والأوعية الدموية ويقترحون مزيدًا من البحث لتحسين النموذج والتحقق من فعاليته عبر مجموعات سكانية متنوعة.

مقدمة

تُعتبر أمراض القلب والأوعية الدموية (CVD) السبب الرئيسي للوفيات العالمية، حيث تم الإبلاغ عن 19.4 مليون حالة وفاة في عام 2021، وتشير التوقعات إلى زيادة العدد إلى ما يقرب من 24 مليون بحلول عام 2030، وفقًا لتقرير العبء العالمي للأمراض (GBD) لعام 2021. تؤكد الطبيعة غير العرضية لـ CVD على الحاجة الملحة لتحديد المخاطر مبكرًا. تؤكد إرشادات ACC/AHA لعام 2010 على أهمية أدوات التنبؤ بالمخاطر، مثل درجة مخاطر القلب والأوعية الدموية في فرامينغهام (FGCRS)، SCORE2، وQRISK2، التي تستخدم عوامل الخطر التقليدية مثل العمر، والجنس، ومستويات الدهون. ومع ذلك، قد تفوت هذه النماذج عوامل خطر إضافية وغالبًا ما تكون مقيدة بأساليب إحصائية كلاسيكية تفترض توزيعات وعلاقات بيانات محددة.

استجابةً لهذه القيود، تستخدم هذه الدراسة تقنيات التعلم الآلي (ML) لتطوير والتحقق من نموذج شامل للتنبؤ بمخاطر CVD باستخدام بيانات من مجموعة بيانات UK Biobank. تقوم الدراسة ببناء نماذج منفصلة لمختلف الحالات القلبية الوعائية، بما في ذلك CVD بشكل عام، وأمراض الأوعية الدموية الدماغية، والأمراض التخثرية. من خلال دمج الدرجات الجينية، والمعايير السريرية، وعوامل نمط الحياة، تهدف الدراسة إلى تعزيز دقة وتطبيق تقييمات مخاطر CVD، مما يسهل في النهاية تقييم المخاطر الفردية في البيئات السريرية.

الطرق

يستعرض قسم “الطرق” تصميم التجربة والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث قاموا بإجراء تحليلات إحصائية لتقييم العلاقات بين المتغيرات. شملت جمع البيانات استبيانًا منظمًا تم إدارته لعينة سكانية، مما يضمن تمثيل ديموغرافي. تضمن الاستبيان أدوات موثوقة لقياس المفاهيم الرئيسية، مما يعزز موثوقية النتائج.

بالإضافة إلى ذلك، استخدمت التحليلات نماذج الانحدار لتقييم تأثير المتغيرات المستقلة على النتائج التابعة. تأكد الباحثون من قوة نتائجهم من خلال اختبارات تشخيصية متنوعة، بما في ذلك فحوصات التعدد الخطي والتباين غير المتجانس. بشكل عام، تم تصميم الإطار المنهجي لتوفير فهم شامل للظواهر قيد التحقيق، مما يسهل تفسير النتائج في سياق الأدبيات الموجودة.

النتائج

يقدم قسم “النتائج” من الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يوضح النتائج الناتجة عن اختبارات متنوعة، مع تسليط الضوء على العلاقات الإحصائية الهامة والاتجاهات الملحوظة في البيانات. غالبًا ما تكون النتائج مصحوبة بأشكال وجداول ذات صلة توضح النتائج بوضوح، مما يسمح بتفسير بصري للبيانات.

في هذا القسم، قد يناقش المؤلفون أيضًا تداعيات نتائجهم فيما يتعلق بالفرضيات المطروحة سابقًا في الدراسة. يقدمون فحصًا شاملاً للنتائج، بما في ذلك أي نتائج غير متوقعة وتفسيراتها المحتملة. بشكل عام، يخدم هذا القسم لتأكيد أهداف البحث ويساهم في الفهم الأوسع للموضوع قيد التحقيق.

المناقشة

تستفيد هذه الدراسة من بيانات UK Biobank، التي تشمل أكثر من 500,000 فرد، لتطوير نموذج تعلم آلي (ML) للتنبؤ بمخاطر أمراض القلب والأوعية الدموية (CVD). كانت النقطة النهائية الرئيسية هي تشخيص CVD، الذي تم تحديده من خلال رموز ICD-10 محددة. أسفرت عملية اختيار صارمة عن 305 متغيرات ذات صلة سريرية، بما في ذلك البيانات الديموغرافية، وعوامل نمط الحياة، والقياسات البدنية، والاختبارات البيولوجية. استخدم النموذج خوارزميات ML متنوعة، بما في ذلك الانحدار Ridge، والانحدار اللوجستي، وغابة عشوائية، واستفاد من تصميم حالة تحكم متداخلة لضمان تقييم قوي. حقق النموذج النهائي، الذي يدمج البيانات السريرية، وعوامل نمط الحياة، والبيانات الجينية، منطقة تحت المنحنى (AUC) تبلغ 0.85، متفوقًا بشكل كبير على أدوات التقييم التقليدية مثل درجة فرامينغهام.

كشفت التحليلات عن أداء تنبؤي متغير عبر أنواع فرعية مختلفة من CVD، حيث حقق النموذج المتكامل AUC قدره 0.73 لأمراض الأوعية الدموية الدماغية و0.66 للأمراض التخثرية، بينما أظهر قوة تنبؤية قوية (AUC قدره 0.83) لأمراض القلب والأوعية الدموية الأخرى. شملت المتنبئات الرئيسية العمر، وضغط الدم الانقباضي، ومستويات السيستاتين C، مما يبرز أهمية كل من عوامل الخطر المعروفة والجديدة. تسلط الدراسة الضوء على الآثار العملية للنموذج في البيئات السريرية، داعيةً إلى استخدامه في التعرف المبكر واستراتيجيات العلاج الشخصية للأفراد ذوي المخاطر العالية. يجب أن تركز الأبحاث المستقبلية على تعزيز دقة النموذج من خلال دمج متنبئات إضافية والتحقق من قابليته للتطبيق عبر مجموعات سكانية متنوعة.

القيود

تعترف الدراسة بعدة قيود قد تؤثر على النتائج. أولاً، على الرغم من أن الدراسة استخدمت نهج تطوير نموذج خالٍ من التسرب مع مجموعة اختبار محجوزة صارمة، إلا أن الحاجة إلى التحقق الخارجي في مجموعات مستقلة تظل ضرورية لتأكيد قابلية نقل النموذج وتطبيقه في السيناريوهات الواقعية. ثانيًا، يؤدي الاعتماد على بيانات من UK Biobank إلى إدخال تحيزات محتملة في الاختيار والمشاركة، مما قد يحد من تعميم النتائج على مجموعات سكانية سريرية أكثر تنوعًا أو ذات مخاطر أعلى.

بالإضافة إلى ذلك، كانت النتائج مستمدة أساسًا من الترميز الإداري والسجلات، مما يثير مخاوف بشأن احتمال التصنيف الخاطئ والتباين في ممارسات الترميز. قد يقلل هذا من أداء النموذج التمييزي، خاصةً بالنسبة لنقاط النهاية الأكثر تباينًا، مثل الأحداث التخثرية. أخيرًا، على الرغم من أن استراتيجية أخذ العينات من حالة التحكم المتداخلة تعزز الكفاءة الحسابية أثناء تطوير النموذج، إلا أن التقييم الإضافي لتقويم المخاطر المطلقة وتحديد العتبات السريرية ضروري، خاصةً ضمن المجموعة الكاملة وعبر مجموعات بيانات خارجية.

Journal: Open Heart, Volume: 13, Issue: 1
DOI: https://doi.org/10.1136/openhrt-2025-003827
PMID: https://pubmed.ncbi.nlm.nih.gov/41720602
Publication Date: 2026-01-01
Author(s): Qingmiao Yang et al.
Primary Topic: Artificial Intelligence in Healthcare

Overview

This research paper addresses the persistent challenge of accurately predicting cardiovascular diseases (CVD), which are the leading cause of global mortality. Traditional risk prediction models, such as the Framingham Cardiovascular Risk Score, often underperform due to their reliance on limited variables and classical statistical methods. To overcome these limitations, the study employs machine learning (ML) techniques to develop a comprehensive CVD risk prediction model using data from the UK Biobank cohort, which includes genetic scores, clinical parameters, and lifestyle factors. The analysis involved 240,644 participants and utilized various ML algorithms, achieving an area under the curve (AUC) of 0.85 for the integrated model, significantly outperforming models based on individual data domains.

Key findings indicate that the integrated model not only enhances predictive accuracy but also identifies critical predictors such as age, systolic blood pressure, cystatin C levels, and waist circumference. The study highlights a 0.12 improvement in AUC over the traditional Framingham score, emphasizing the model’s potential for clinical application in early risk identification and personalized interventions. Additionally, a streamlined version of the model, utilizing only 20 easily obtainable predictors, maintains high predictive performance, making it practical for clinical use. The authors advocate for the integration of ML-based risk prediction tools into cardiovascular prevention strategies and suggest further research to refine the model and validate its effectiveness across diverse populations.

Introduction

Cardiovascular diseases (CVD) are the leading cause of global mortality, with 19.4 million deaths reported in 2021 and projections suggesting an increase to nearly 24 million by 2030, as per the Global Burden of Disease (GBD) 2021 report. The asymptomatic nature of CVD emphasizes the urgent need for early risk identification. The 2010 ACC/AHA guidelines underscore the importance of risk prediction tools, such as the Framingham Cardiovascular Risk Score (FGCRS), SCORE2, and QRISK2, which utilize traditional risk factors like age, gender, and lipid levels. However, these models may miss additional risk factors and are often constrained by classical statistical methods that assume specific data distributions and relationships.

In response to these limitations, this study employs machine learning (ML) techniques to develop and validate a comprehensive CVD risk prediction model using data from the UK Biobank cohort. The research constructs separate models for various cardiovascular conditions, including overall CVD, cerebrovascular diseases, and thrombotic diseases. By incorporating genetic scores, clinical parameters, and lifestyle factors, the study aims to enhance the precision and practical application of CVD risk assessments, ultimately facilitating individualized risk evaluation in clinical settings.

Methods

The “Methods” section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, employing statistical analyses to assess the relationships between variables. Data collection involved a structured survey administered to a sample population, ensuring a representative demographic. The survey included validated instruments to measure key constructs, enhancing the reliability of the findings.

Additionally, the analysis employed regression models to evaluate the impact of independent variables on the dependent outcomes. The researchers ensured the robustness of their results through various diagnostic tests, including checks for multicollinearity and heteroscedasticity. Overall, the methodological framework was designed to provide a comprehensive understanding of the phenomena under investigation, facilitating the interpretation of the results in the context of existing literature.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It details the outcomes of various tests, highlighting significant statistical relationships and trends observed in the data. The results are often accompanied by relevant figures and tables that illustrate the findings clearly, allowing for a visual interpretation of the data.

In this section, the authors may also discuss the implications of their findings in relation to the hypotheses posed earlier in the study. They provide a thorough examination of the results, including any unexpected outcomes and their potential explanations. Overall, this section serves to substantiate the research objectives and contributes to the broader understanding of the topic under investigation.

Discussion

This study leverages data from the UK Biobank, encompassing over 500,000 individuals, to develop a machine learning (ML) model for predicting cardiovascular disease (CVD) risk. The primary endpoint was the diagnosis of CVD, identified through specific ICD-10 codes. A rigorous selection process yielded 305 clinically relevant variables, including demographics, lifestyle factors, physical measurements, and biological assays. The model employed various ML algorithms, including Ridge Regression, Logistic Regression, and Random Forest, and utilized a nested case-control design to ensure robust evaluation. The final model, integrating clinical, lifestyle, and genetic data, achieved an area under the curve (AUC) of 0.85, significantly outperforming traditional risk assessment tools like the Framingham score.

The analysis revealed variable predictive performance across different CVD subtypes, with the integrated model achieving an AUC of 0.73 for cerebrovascular diseases and 0.66 for thrombotic diseases, while demonstrating strong predictive power (AUC of 0.83) for other CVDs. Key predictors included age, systolic blood pressure, and cystatin C levels, emphasizing the importance of both established and novel risk factors. The study highlights the model’s practical implications for clinical settings, advocating for its use in early identification and personalized treatment strategies for high-risk individuals. Future research should focus on enhancing model accuracy through the integration of additional predictors and validating its applicability across diverse populations.

Limitations

The research acknowledges several limitations that may impact the findings. Firstly, while the study employed a leakage-free model development approach with a strict held-out test set, the need for external validation in independent cohorts remains crucial to confirm the model’s transportability and its applicability in real-world scenarios. Secondly, the reliance on data from the UK Biobank introduces potential selection and participation biases, which could limit the generalizability of the results to more diverse or higher-risk clinical populations.

Additionally, the outcomes were primarily derived from administrative and registry coding, which raises concerns about potential misclassification and variability in coding practices. This could diminish the model’s discriminative performance, especially for more heterogeneous endpoints, such as thrombotic events. Lastly, although the nested case-control sampling strategy enhances computational efficiency during model development, further assessment of absolute risk calibration and clinical thresholding is necessary, particularly within the full cohort and across external datasets.

شارك: