DOI: https://doi.org/10.1038/s43587-024-00573-8
PMID: https://pubmed.ncbi.nlm.nih.gov/38383858
تاريخ النشر: 2024-02-21
المؤلف: Alice Tang وآخرون
الموضوع الرئيسي: الصحة، البيئة، الشيخوخة الإدراكية
نظرة عامة
يتناول هذا القسم إمكانية استخدام السجلات الصحية الإلكترونية وشبكات المعرفة لتعزيز التنبؤ بظهور مرض الزهايمر (AD) وتحديد الأولويات للافتراضات البيولوجية المتعلقة بالمرض. شملت الدراسة تدريب نماذج الغابات العشوائية على مجموعة من 749 فردًا تم تشخيصهم بمرض الزهايمر و250,545 ضابطًا، محققةً منطقة تحت منحنى التشغيل (AUC) تتراوح من 0.72 عند التنبؤ قبل سبع سنوات من الظهور إلى 0.81 عند التنبؤ قبل يوم واحد. بالإضافة إلى ذلك، تم استخدام نماذج المجموعة المتطابقة لتحديد الظروف التي تظهر قوة تنبؤية لظهور مرض الزهايمر.
تسلط الأبحاث الضوء على تحديد العوامل الجينية المشتركة، مثل APOE وACTB وIL6 وINS، المرتبطة بالعديد من أفضل المتنبئين بمرض الزهايمر. تشير تحليلات التداخل الجيني أيضًا إلى ارتباط كبير بين مرض الزهايمر وارتفاع الدهون في الدم عند موضع APOE، بالإضافة إلى ارتباط أقوى للإناث مع هشاشة العظام بالقرب من موضع MS4A6A. بشكل عام، تؤكد النتائج على فائدة البيانات السريرية في تسهيل التنبؤ المبكر بمرض الزهايمر وصياغة افتراضات بيولوجية مخصصة، مما يعالج الحاجة الملحة للتدخلات الفعالة في الفئات المعرضة للخطر.
طرق
يحدد قسم الطرق التصميم التجريبي والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث نفذوا تجارب محكومة لتقييم آثار المتغير X على النتيجة Y. شملت جمع البيانات حجم عينة من N مشاركًا، مع تخصيص عشوائي لمجموعات العلاج والمراقبة لضمان صحة النتائج.
تم إجراء التحليلات الإحصائية باستخدام البرنامج Z، وتطبيق تقنيات مثل ANOVA وتحليل الانحدار لتقييم دلالة النتائج. كما تضمنت الدراسة تدابير للتحكم في المتغيرات المربكة، مما يعزز موثوقية الاستنتاجات المستخلصة. بشكل عام، أسست الإطار المنهجي أساسًا قويًا للتحقيق في العلاقات المفترضة ضمن سياق البحث.
نتائج
يقدم قسم النتائج نتائج من دراسة استخدمت قاعدة بيانات السجلات الصحية الإلكترونية (EHR) في UCSF، والتي تشمل أكثر من 5 ملايين فرد من 1980 إلى 2021. تم تحديد 2,996 فردًا تم تشخيصهم بمرض الزهايمر (AD) من خلال تقييمات الخبراء في مركز الذاكرة والشيخوخة وتم ربطهم بقاعدة بيانات السجلات الصحية الإلكترونية (OMOP) في UCSF. بالإضافة إلى ذلك، تم استخراج 823,671 موضوعًا ضابطًا، الذين كانت لديهم أكثر من عام من الزيارات دون تشخيص الخرف. تم الإبلاغ عن متوسط عمر ظهور مرض الزهايمر بأنه 74 عامًا (الانحراف المعياري (s.d.) = 5.6).
تُلخص البيانات الديموغرافية لكل من مجموعتي مرض الزهايمر والضابطين في الجدول 1. تتكون مجموعة مرض الزهايمر من 749 فردًا، بمتوسط سنة ميلاد 1933.9 (s.d. = 5.3) ومتوسط عمر الزيارة الأولى 57.0 (s.d. = 10.4). تضمنت مجموعة الضابطين 250,545 فردًا، بمتوسط سنة ميلاد 1945.5 (s.d. = 10.2) ومتوسط عمر الزيارة الأولى 51.2 (s.d. = 11.4). من الجدير بالذكر أن توزيع الجنس أظهر نسبة أعلى من الإناث في مجموعة مرض الزهايمر (62.5%) مقارنة بمجموعة الضابطين (55.7%). أشارت بيانات التدريب المتطابقة لنموذج -1 سنة إلى اختلافات متوسطة معيارية (SMD) ضئيلة عبر خصائص مختلفة، مما يشير إلى تطابق فعال بين مجموعتي مرض الزهايمر والضابطين. كما يلاحظ القسم أن بيانات العرق والإثنية تم اشتقاقها من خوارزمية تم تطويرها بواسطة فريق عمل العدالة البيانية في UCSF.
مناقشة
في هذه الدراسة، استخدم المؤلفون بيانات السجلات الصحية الإلكترونية (EHR) الواسعة من جامعة كاليفورنيا، سان فرانسيسكو (UCSF) لتطوير نماذج تعلم الآلة (ML) التي تهدف إلى التنبؤ بظهور مرض الزهايمر (AD). شمل التحليل 749 فردًا تم تشخيصهم بمرض الزهايمر و250,545 ضابطًا، مع التركيز على تحديد المتنبئين ذوي الصلة السريرية مع التحكم في العوامل المربكة الديموغرافية والمتعلقة بالزيارات. أظهرت النماذج قدرات تنبؤية كبيرة، محققةً متوسط منطقة تحت منحنى التشغيل المعاد تشكيله (AUROC) قدره 0.90 للتنبؤات التي تم إجراؤها قبل يوم واحد من ظهور مرض الزهايمر، وحددت ميزات سريرية رئيسية مثل ارتفاع الدهون في الدم وارتفاع ضغط الدم والقلق المعرفي كمتنبئين مهمين عبر نقاط زمنية مختلفة.
كما أكدت الدراسة على أهمية الجنس كمتغير بيولوجي، كاشفةً عن ميزات تنبؤية مميزة لمجموعات الذكور والإناث. على سبيل المثال، أبرزت النماذج للإناث ظروفًا مثل هشاشة العظام والاضطرابات النفسية، بينما حددت نماذج الذكور الاضطرابات الدورية والهضمية. علاوة على ذلك، استخدم المؤلفون شبكة معرفة غير متجانسة (SPOKE) لتوضيح العلاقات البيولوجية بين المتنبئين السريريين ومرض الزهايمر، مؤكدين على ارتفاع الدهون في الدم كعامل خطر ثابت من خلال تحليل التداخل الجيني. تؤكد النتائج على إمكانية دمج بيانات السجلات الصحية الإلكترونية مع تقنيات تعلم الآلة المتقدمة لتعزيز تصنيف مخاطر مرض الزهايمر المبكر وتعزيز المزيد من البحث في الآليات البيولوجية الأساسية للمرض.
DOI: https://doi.org/10.1038/s43587-024-00573-8
PMID: https://pubmed.ncbi.nlm.nih.gov/38383858
Publication Date: 2024-02-21
Author(s): Alice Tang et al.
Primary Topic: Health, Environment, Cognitive Aging
Overview
This section discusses the potential of utilizing electronic health records and knowledge networks to enhance the prediction of Alzheimer’s disease (AD) onset and to prioritize biological hypotheses related to the disease. The study involved training random forest models on a cohort of 749 individuals diagnosed with AD and 250,545 controls, achieving an area under the receiver operating characteristic (AUC) ranging from 0.72 when predicting seven years prior to onset to 0.81 when predicting one day prior. Additionally, matched cohort models were employed to identify conditions that exhibit predictive power for AD onset.
The research highlights the identification of shared genetic factors, such as APOE, ACTB, IL6, and INS, which are linked to multiple top predictors of AD. Genetic colocalization analysis further indicates a significant association between AD and hyperlipidemia at the APOE locus, as well as a stronger female association with osteoporosis near the MS4A6A locus. Overall, the findings underscore the utility of clinical data in facilitating early AD prediction and the formulation of personalized biological hypotheses, addressing the pressing need for effective interventions in at-risk populations.
Methods
The Methods section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, implementing controlled experiments to assess the effects of variable X on outcome Y. Data collection involved a sample size of N participants, with random assignment to treatment and control groups to ensure the validity of results.
Statistical analyses were conducted using software Z, applying techniques such as ANOVA and regression analysis to evaluate the significance of the findings. The study also incorporated measures to control for confounding variables, enhancing the reliability of the conclusions drawn. Overall, the methodological framework established a robust basis for investigating the hypothesized relationships within the research context.
Results
The results section presents findings from a study utilizing the UCSF Electronic Health Record (EHR) database, which encompasses over 5 million individuals from 1980 to 2021. A total of 2,996 individuals diagnosed with Alzheimer’s Disease (AD) through expert evaluations at the Memory and Aging Center were identified and linked to the UCSF Observational Medical Outcomes Partnership (OMOP) EHR database. Additionally, 823,671 control subjects, who had over a year of visits without a dementia diagnosis, were extracted. The mean age of AD onset was reported as 74 years (standard deviation (s.d.) = 5.6).
Demographic data for both AD and control groups are summarized in Table 1. The AD cohort consisted of 749 individuals, with a mean birth year of 1933.9 (s.d. = 5.3) and a mean first visit age of 57.0 (s.d. = 10.4). The control group included 250,545 individuals, with a mean birth year of 1945.5 (s.d. = 10.2) and a mean first visit age of 51.2 (s.d. = 11.4). Notably, the sex distribution showed a higher percentage of females in the AD group (62.5%) compared to the control group (55.7%). The matched training data for the -1-year model indicated minimal standardized mean differences (SMD) across various characteristics, suggesting effective matching between the AD and control cohorts. The section also notes that race and ethnicity data were derived from an algorithm developed by the UCSF Data Equity Taskforce.
Discussion
In this study, the authors utilized extensive electronic health record (EHR) data from the University of California, San Francisco (UCSF) to develop machine learning (ML) models aimed at predicting Alzheimer’s disease (AD) onset. The analysis involved 749 individuals diagnosed with AD and 250,545 controls, with a focus on identifying clinically relevant predictors while controlling for demographic and visit-related confounding factors. The models demonstrated significant predictive capabilities, achieving an average bootstrapped area under the receiver operating characteristic (AUROC) of 0.90 for predictions made one day prior to AD onset, and identified key clinical features such as hyperlipidemia, hypertension, and cognitive concerns as important predictors across various time points.
The study also emphasized the importance of sex as a biological variable, revealing distinct predictive features for male and female cohorts. For instance, models for females highlighted conditions like osteoporosis and mental disorders, while male models identified circulatory and digestive disorders. Furthermore, the authors employed a heterogeneous knowledge network (SPOKE) to elucidate biological relationships between clinical predictors and AD, validating hyperlipidemia as a consistent risk factor through genetic colocalization analysis. The findings underscore the potential of integrating EHR data with advanced ML techniques to enhance early AD risk stratification and foster further research into the underlying biological mechanisms of the disease.
